STALE MEASUREMENTPast this project's own 2-release window: the published result was measured with v0.32.0, 9 releases ago. Why, and what unblocks it

ProductEvidenceTop 10LeaderboardCompliancePricingDocsStar on GitHub Quickstart
0 OF 3 PUBLISHED · 3 OF 3 SPOTS OPEN · AS OF 20 AUGUST 2026

No case studies yet. Here is exactly what one will contain.

Zero engagements have completed, so there is nothing here to read about a customer — and a page that padded that gap with an illustrative composite would be the first fabricated thing on this site. What follows instead is the contract: what a published study contains, what anonymisation removes, what happens when the result is bad, and a worked example built from Provael's own measured run.

The shape

What a published study contains

The same shape as every measured write-up already on /findings, because a case study that reported differently from the research would be marketing wearing the research's clothes:

  • The headline attack-success rate, with its 95% Wilson confidence interval.
  • The benign false-positive control the rate is read against.
  • The attack families that did not transfer, reported at the same weight.
  • A "what this is not" section stating the limits in the author's own words.
  • The transfer status: measured on a real policy, or stub-validated scaffolding.

It is indexed, linked from the research feed, and included in this site's Markdown twins so an agent evaluating vendors can read it without rendering the page.

The contract

What anonymisation removes — and what it cannot

A design partner may publish co-branded or anonymised. Anonymised is not a softer version of the result; it is the same measurement with the identifying detail stripped. A study becomes "a humanoid manipulation policy" rather than a named product.

Removed

  • Your organisation's name and logo.
  • The checkpoint identifier.
  • Any model or product name.
  • The task descriptions, if they are proprietary.
  • Any trace detail that would identify the system.

Retained

  • The attack families that were run.
  • Every rate, with its 95% Wilson confidence interval.
  • The benign false-positive control.
  • The families that did not transfer — the nulls.
  • The caveats: simulation only, stated sample size, stated limits.

The one thing anonymisation cannot do is remove the result itself. If a finding is only publishable when nobody can tell it happened at all, this is the wrong programme and the standard rate is the right one. That sentence is on /design-partners too, and it is load-bearing in both places.

The objection

What happens if the result is bad

Then we publish it — and that is the answer given before it is asked.Why we publish null results was written before the design-partner programme existed. A red team that only reports when it finds something has no denominator, and a vendor who suppresses its own bad results has told you exactly what its good ones are worth.

Two different things hide under "bad result", and they are worth separating before a signature:

  • A high ASR is a finding about the policy. Published with the remediation and the free retest, it is a stronger story than a clean sheet — it is the shape most of the useful write-ups in this field take.
  • A null result — the attacks did not transfer — is a publishable measurement too. The reference run below already reports honest 0% nulls for 2 of its 3 adversarial families.
Worked example

Reference case: Provael's own run — not a customer engagement

Read this label first. The numbers below are Provael's own published result, already on /results and in the product repo's committed report.json. No customer was involved and no engagement produced them. They are here to show the shape and rigour a study will take, not to imply a client relationship that does not exist.

Policy smolvla on LIBERO, task libero_object (all 10 tasks), 5 seeds, 350 episodes in total.

AttackFamilyTrialsRate95% CI
nonebaseline2/504%1-13%
roleplayinstruction44/5088%72-100%
goal_substitutioninstruction15/5030%19-44%
paraphraseinstruction3/506%2-16%
patchvisual0/500%0-7%
decoy_objectvisual0/500%0-7%
scene_textinjection0/500%0-7%
mcp_tool_descinjectionN/AN/AN/A

Instruction family aggregate: 62/150 = 41.3% [34-49%]. Benign control: 2/50 = 4% — the false-positive floor every rate above is read against.

Two of the three adversarial families measured zero, and they are in the table at the same weight as the one that transferred. mcp_tool_desc reads N/A, not 0% — it had no surface in this suite, and an attack that could not run is excluded from the denominator rather than scored as a pass.

Next

Becoming the first published study

The founding cohort is the mechanism that puts a real engagement on this page. 3 of 3 spots are open as of 20 August 2026, at a founding rate traded for permission to publish.

The design-partner programme →Book a scoping call