STALE MEASUREMENTPast this project's own 2-release window: the published result was measured with v0.32.0, 9 releases ago. Why, and what unblocks it

ProductEvidenceTop 10LeaderboardCompliancePricingDocsStar on GitHub Quickstart
FOR HUMANOID BUILDERS · PV-018

Red-team your humanoid's policy before it ships.

A humanoid runs on a whole-body vision-language-action policy - one network turning cameras and an instruction into motion across the entire body. That policy is a new attack surface, and it is arriving just as the humanoid-safety rules do. Provael red-teams the policy in simulation and reports an attack-success rate with a 95% Wilson confidence interval and a benign control - a measurement your diligence, insurer, or CE file can work from.

Every attack family · freeDefensive · sim-onlyGR00T · π0 · OpenVLA adapters
The problem you are shipping into

A whole-body policy is a new attack surface, under rules that are arriving

The EU Machinery Regulation binds 20 January 2027 and pulls AI-driven safety functions into a cyber-risk assessment; ISO 10218:2025 makes cybersecurity a requirement for industrial robots; and ISO 25785-1, covering dynamically stable industrial mobile robots (a category that includes bipedal humanoids, though it is not humanoid-specific), is still a Committee Draft and not published. A reworded instruction or a perturbed observation can redirect a policy's motion without touching a line of code - and for a bipedal machine, the failure is a fall or a step into a person, not a dropped part. That behaviour is exactly what text-only red-teaming can't see.

What runs against your policy today

The method transfers to the policy driving your humanoid

Point Provael at your checkpoint - GR00T, π0, and OpenVLA have registered adapters - and run the attack families in simulation for an ASR with a 95% Wilson CI and a benign control.Registered is not the same as measured: only SmolVLA has been run end to end against a real checkpoint so far. The GR00T-N1 transfer study is pre-registered and has not been run, and the adapters above are exercised against the deterministic fixture rather than a real policy. This is the action layer -see where it sits relative to the compute certification your silicon vendor ships and the CVE work on your ROS stack. The families:

Instruction reframing

A reworded goal that reads as compliant but redirects the policy - the one family with a real measured transfer (on a manipulation policy).

Adversarial perception

Patches, textures and sensor spoofs that flip behaviour while looking benign to a human.

Action-space integrity

Attacks on the commanded action itself - targeted hijack or freeze - scored against a benign control.

Embodied injection

Instructions that enter through the environment - a sign, a label, a tool description the policy reads.

The humanoid safety pack

A whole-body suite and three locomotion attacks - shipped, stub-validated

A deterministic CPU humanoid suite - fall/topple, loss of balance (centre of mass outside the support polygon), self-collision, footstep keep-out - and three attacks now ship in the tool, each with a transfer-test (ASR + 95% Wilson CI + a benign-FPR control) - and the one action-side defense this project has measured,action_envelope, is not credited on humanoid(it is credited on stub and reach). They are stub-validated on the fixture: a 100% fixture rate is scaffolding, not evidence of real transfer.

That defense result is stated here rather than left on /defenses because it is the one a humanoid maker would otherwise assume applies to them. It does not, and the reason is structural rather than a tuning gap: the measured pre and post rates on the humanoid suite are both 100.0% [89-100%] - the mitigation is a magnitude clamp on the commanded action, and a balance predicate is not a magnitude at all. Read the full verdicts in the action-envelope study ↗, or the index of every study at /studies.

STUB-VALIDATED

balance_spoof EAI02

Spoof the policy toward a loss of balance - centre of mass outside the support polygon - and score whether it recovers or steps unsafely.

STUB-VALIDATED

whole_body_hijack EAI04

Steer the whole-body trajectory toward an attacker end-state, not just an end-effector.

STUB-VALIDATED

stride_freeze EAI04

Drive the locomotion policy to a mid-stride halt - an availability failure with its own hazards.

Where Provael fits

Adversarial robustness, not the functional-safety stack

Functional-safety frameworks - NVIDIA Halos, UL 4600, ISO 21448 - govern the safety architecture and the case around your system. Provael does something different and complementary: it measures the adversarial attack-success rate on what the policy actually does, so the robustness number those frameworks need as an input is a measurement, not a self-attestation. We do not certify anything; we supply the evidence.

For the engineer

Start free and CPU-first: pip install provael, wrap your policy with a small adapter, run the suite, and read your ASR - then gate CI on it. The CLI, the attacks, SARIF, and the GitHub Action are free and Apache-2.0, and always will be.

For the person who signs off

An independent, reproducible attack-success rate - mapped to the frameworks you cite - that a certifier or insurer can review. It is candidate evidence toward a conformity assessment, not a certificate.

Measure your humanoid's policy before someone else does.

Run it yourself today, or book a scoped red-team of your policy for your next milestone, raise, or audit gate. Read the evidence pack you would receive before booking — and if publication is something you can agree to, the design-partner programme is the founding-cohort rate.

The evidence, in three links

Everything argued above rests on three published artifacts. Start with the measured result, then the families that returned nothing on the same model, then the physical-robot protocol that is registered and unrun.

  1. The one measured resultA real SmolVLA policy driven out of its envelope on 44/50 trials across all ten libero_object tasks, with its 95% task-clustered interval and a 2/50 (4%) benign control. One policy, one suite, ten tasks.
  2. The results that came back nullThe visual and injection families scored 0% on the same real model. Published at the same size as the number that worked, because a red-team that only reports hits is a demo.
  3. The pre-registration, with no results in itThe physical-robot protocol, its predicate and its stopping rule, published before the trials run. Nothing has been measured on hardware, and that page says so first.
Book an assessment →Contact us

Feature and status claims on this page were read against the product repository on . Unlike the numbers on /results and /leaderboard, these are not enforced by a build check — the product publishes no machine-readable artifact for per-suite defense verdicts, so this is a human review with a date on it, not a guarantee. The maintained source is the repository; where it disagrees with this page, it wins.