STALE MEASUREMENTPast this project's own 2-release window: the published result was measured with v0.32.0, 9 releases ago. Why, and what unblocks it

ProductEvidenceTop 10LeaderboardCompliancePricingDocsStar on GitHub Quickstart
FOR HUMANOID & ROBOTICS TEAMS · PV-016

Ship the robot. Not the vulnerability.

You’re months from a warehouse, a home, or a public demo floor. Your policy’s safety is about to become your next investor’s diligence question, your first enterprise customer’s security questionnaire, and your insurer’s renewal condition. Provael gives you independent, reproducible evidence that your robot’s policy was tested - before a regulator, an insurer, or a lawsuit asks.

Free CLI · pip install provaelCPU-first · CI-gatedAttack-success rate + 95% CI + benign controlBuilding a humanoid? →
The stakes

The action layer is where the real incidents already are

2026 has already produced two CVSS-9.8 vulnerabilities on the exact stack robotics teams run: an unauthenticated RCE in Hugging Face LeRobot (the default open VLA stack), and a command injection in Universal Robots controllers - both let an attacker move a physical machine. And a reworded instruction can hijack a vision-language-action policy’s motion without touching a single line of code. That’s the layer text-only red-teaming can’t see. See the incident tracker →

What you get

A defensible safety story, on demand

  • An attack-success rate for your policy - with a 95% Wilson confidence interval and a benign-control false-positive rate, not a one-off demo.
  • A CI gate (SARIF + GitHub Action) that fails the build if robustness regresses - so safety is enforced, not assumed.
  • Coverage mapped to the Embodied AI Security Top 10 - the vocabulary your diligence, customers, and auditors can follow.
  • An evidence pack you can file when a raise, a customer, or an insurer needs more than self-attestation.
The attack surface

Four ways a policy gets redirected — and which one we have actually measured

A vision-language-action policy turns language and perception into robot actions, so the attack surface is the action, not the answer. These are the four surfaces a startup ships with. Each links to its full entry in the Embodied AI Security Top 10.

How it fits

Start free, in an afternoon

pip install provael, run the suite against the CPU stub or a real policy you wrap with a tiny adapter, and read your ASR. No GPU required to start. When you need an independent, signed result for a specific milestone, book a scoped assessment - see pricing.

Prove it before someone else does.

Run it yourself today, or book an independent red-team assessment for your next raise or deployment. See the evidence pack you would receive first, and — if you can agree to publication — the founding-cohort design-partner rate.

The evidence, in three links

Everything argued above rests on three published artifacts. Start with the measured result, then read what did not work, then check that the hardware claim we have not yet made was designed before anyone could see the data.

  1. The one measured resultA real SmolVLA policy driven out of its envelope on 44/50 trials across all ten libero_object tasks, with its 95% task-clustered interval and a 2/50 (4%) benign control. One policy, one suite, ten tasks.
  2. The results that came back nullThe visual and injection families scored 0% on the same real model. Published at the same size as the number that worked, because a red-team that only reports hits is a demo.
  3. The pre-registration, with no results in itThe physical-robot protocol, its predicate and its stopping rule, published before the trials run. Nothing has been measured on hardware, and that page says so first.
Quickstart →Book an assessment

Feature and status claims on this page were read against the product repository on . Unlike the numbers on /results and /leaderboard, these are not enforced by a build check — the product publishes no machine-readable artifact for per-suite defense verdicts, so this is a human review with a date on it, not a guarantee. The maintained source is the repository; where it disagrees with this page, it wins.