The robots are shipping. The standards aren’t ready. Here’s evidence you can work from.
The EU Machinery Regulation binds on 20 January 2027 and pulls safety components with self-evolving, machine-learned behaviour into mandatory third-party conformity assessment - yet dedicated humanoid / robot-AI standards don’t exist yet, and assessment capacity is still ramping. That is not a deadline, it is an 18-month evidence gap: the delegated acts carrying the AI robustness requirements into Annex III of that Regulation do not apply until 2 August 2028. Provael gives you a standardized, measurable methodology to fill that gap: an attack-success rate you can review, compare, and underwrite. If you are mapping who covers what, here is the layer map of the physical-AI safety stack - what Halos, UL 3300 and the CVE work each certify, and where the learned policy falls between them.
You’re being asked to assess systems with no settled test method
A learned robot policy is probabilistic, and a single “task success” number hides the two failures that matter - a hijacked trajectory and a frozen actuator. Without a reproducible measurement, there is nothing objective to review or price. Provael turns “is this policy safe enough?” into a number with a stated confidence interval and a benign control - the kind of measurable performance data an assessment or a parametric underwriting model can actually use.
Independent, comparable, defensible
- Attack-success rate with a 95% Wilson CI and a matched benign-FPR control - a defensible measurement, not a vendor’s self-attestation. To date, one measured result: the roleplay attack put SmolVLA out of its keep-out envelope on 44/50 trials [72-100% CI] vs 4% on the benign control - a keep-out violation, not a demonstration of attacker control (full method and caveats).
- Open, tool-ingestible outputs: SARIF, OSCAL assessment-results, a CycloneDX ML-BOM, and a print-to-PDF conformity dossier.
- Mappings to the frameworks you already cite - EU AI Act, EU Machinery Regulation, ISO 10218:2025, IEC 62443, NIST AI RMF.
- The CRA Article 14 clock - what a manufacturer owes in the first 24 hours, and what "we fixed it" has to mean when the exploited component is a learned policy rather than firmware.
- The Korea AI Framework Act crosswalk - the first comprehensive AI statute outside the EU, in force since 22 January 2026, applying to acts abroad that affect the Korean market. Its Article 34(1) duties on risk management, human supervision and safety documentation are the ones a red-team result speaks to; it establishes no Korean conformity or registration position.
- Honest by construction: published null results and transfer statements, so a number never claims more than it measured.
The insurer view, on demand
provael attest --run <report> --profile insurer emits a signed assurance statement over a real ASR run - the headline rate with its 95% Wilson CI and benign-FPR control, plus the honest per-family which-attacks-transfer-on-the-real-model table (each labelled measured-real-transfer or stub-validated). The --profile iso-10218-2 view routes thesame evidence to an IEC 62443 SL2 routing target — SL2 being protection against intentional violation using simple means, and a fixed property of this profile rather than a judgement about what your cell requires — and a cert-readiness cross-reference maps it to the functional- and AI-safety frameworks it is an input to (NVIDIA Halos, UL 4600, ISO 21448, ISO/PAS 8800). Ed25519-signed, verifies offline. Evidence, not certification.
The underwriting gap, in the market’s own words
This page argues a measured attack-success rate is the kind of thing an underwriter can use. That is an argument, not a finding, so here is what is actually on the record. Nothing below says an insurer requires a red-team result, or that a Provael run buys cover or moves a premium. No insurer has reviewed a Provael artifact. And nothing here rests on a robot having hurt anyone: the market is moving on exposure and on the absence of actuarial history, which is checkable, and this project has no injury story to tell.
- A standard generative-AI exclusion now exists for US general liabilityNamed body, specific action
In January 2026 the Insurance Services Office — the US body that drafts standard policy forms, not the International Organization for Standardization — introduced a generative-AI exclusion for commercial general liability. The endorsement excludes bodily injury, property damage, and personal or advertising injury arising out of, or attributable to, generative AI.
What this does not establish: Whether it reaches a vision-language-action robot policy is a question about that endorsement’s wording and your own programme, not something this site can answer. It is evidence that the market is drawing a line around AI-attributable harm, not evidence about where the line falls for you.
Fenwick, "The End of ‘Silent AI’?", 15 June 2026
- A carrier has written a programme specifically for autonomous robotsNamed carrier, reported by trade press
Axis Insurance built a programme for companies that make and deploy autonomous robots, reported as covering bodily injury and property damage from AI navigation or perception failures, physical damage from a cyberattack that takes over a robot’s controls, and production loss from a software update or sensor failure that stops a robot without damaging it.
What this does not establish: The policy wording is not public here. This is trade-press description of a product, not a reading of a filed form, and nothing in it references a red-team measurement.
PYMNTS, "Insurers Face a New Risk From Autonomous Robots", 23 June 2026
- The stated obstacle is missing history, not missing appetiteNamed carrier, reported by trade press
The same reporting gives the underwriting problem in the underwriters’ own terms: "Insurers usually price commercial risk using past claims, equipment records and operating controls. Physical AI gives them less history and more variables." The same robot is a different risk depending on layout, worker interaction and software version, and a flaw in one widely used model is correlated across every business running it.
What this does not establish: That a reproducible adversarial measurement fills that gap. It is an argument this project finds persuasive and has not tested with an underwriter; no insurer has reviewed a Provael artifact, and none is quoted here saying one would help.
PYMNTS, "Insurers Face a New Risk From Autonomous Robots", 23 June 2026
- Red-teaming evidence is reported to be asked for, by insurers nobody has namedUnnamed insurers, secondary report
One trade article states that "other insurers have launched AI security riders in 2026, demanding proof of red-teaming and documented risk assessments before extending cover". It names Chubb as covering certain AI incidents while excluding correlated losses, and names no insurer at all for the red-teaming claim, which it attributes to an advisory firm’s industry commentary.
What this does not establish: Anything, really — and it is carried here precisely because it is the row that would flatter this project most if it were solid. No carrier, no product, no policy wording, no definition of what counts as red-teaming. Do not read it as a requirement, and do not let anyone tell you a Provael run satisfies one.
FinTech Global, "Why autonomous AI could void your cyber insurance in 2026", 28 July 2026
Hand-maintained, checked 6 September 2026. Nothing regenerates it, so it goes stale unless someone looks. “ISO” in the first row is the Insurance Services Office, the US body that drafts standard policy forms — not the International Organization for Standardization, whose ISO 10218 this page cites above.
Evidence, not a certificate
Provael does not certify anything and doesn’t pretend to. It produces the reproducible, adversarial evidence a notified body reviews or an insurer prices - the comparability artifact that’s missing today. You keep the judgment; we supply the measurement.
Talk to us about a reference methodology.
If you assess or underwrite AI-driven robots, let’s map Provael’s evidence to your process - or run it on a real policy together. See the compliance & attestation report on pricing, the evidence pack itself in full, and how a per-checkpoint signed attestation works in CI. Vendors you assess may qualify for the design-partner programme, which publishes its results — /case-studies sets out what a published study contains and what anonymisation removes, and states plainly that none have completed yet.
Before that conversation, the diligence a certifier or underwriter asks for first is already published: the answered security questionnaire, the sub-processor list, the DPA and rules-of-engagement templates, and the insurance and entity position stated as it actually is, all at /trust. Whether anyone independent has reproduced these results — nobody has, yet — is at /verification, with the commands to do it.
The evidence, in three links
Everything argued above rests on three published artifacts — and for your purposes the third matters most. A protocol registered in advance of its own trials is checkable evidence of method; a single measured number is not. Read the pre-registration first, then the measured result and the nulls it sits beside.
- The pre-registration, with no results in itThe physical-robot protocol, its predicate and its stopping rule, published before the trials run. Nothing has been measured on hardware, and that page says so first.
- The one measured resultA real SmolVLA policy driven out of its envelope on 44/50 trials across all ten libero_object tasks, with its 95% task-clustered interval and a 2/50 (4%) benign control. One policy, one suite, ten tasks.
- The results that came back nullThe visual and injection families scored 0% on the same real model. Published at the same size as the number that worked, because a red-team that only reports hits is a demo.
Feature and status claims on this page were read against the product repository on . Unlike the numbers on /results and /leaderboard, these are not enforced by a build check — the product publishes no machine-readable artifact for per-suite defense verdicts, so this is a human review with a date on it, not a guarantee. The maintained source is the repository; where it disagrees with this page, it wins.