STALE MEASUREMENTPast this project's own 2-release window: the published result was measured with v0.32.0, 9 releases ago. Why, and what unblocks it

ProductEvidenceTop 10LeaderboardCompliancePricingDocsStar on GitHub Quickstart
RELEASES · READ FROM THE REPOSITORY

Every release, and how old the measurements are.

Read from the product repository’s CHANGELOG.md when this page was built, not transcribed. The current release is 0.41.2, and the newest entry below is 0.41.2 (9 September 2026).

last measured

0 days ago

Source: watch/freshness.json ↗, which the tool regenerates and which decays on its own rather than being asserted here.

This site

Changes to provael.com

Distinct from the tool’s releases below. A version string on a page can go stale without any release happening, which is how four of them did.

E-2026-11 on /errata, and the regulatory clock says how wide the machinery-standards gap is ·

  • /errata carries E-2026-11: the product’s gradient_patch attack, as released from 0.39.1 to 0.41.2, started its search where its objective’s gradient is exactly zero and so could not move a frame; the PushT figures it shipped with are an external script’s, not the module’s. No page on this site quoted them and no published rate moves.
  • The Machinery Regulation entry of the regulatory clock now records what a 13 September check found: no harmonised standard yet cited under the Regulation, the first list expected in Q4 2026, prEN 50742 (formal vote September, publication planned November 2026) as the one AI-adjacent standard on course for 20 January 2027, and ISO/IEC TS 22440 and IEC 61508 Ed. 3 both landing in 2027. The CRA entry separates the manufacturer’s binding duty from Provael’s own voluntary alignment.

An honesty pass: the retired single-task result is gone from every surface, and a note corrected itself ·

  • The single-task result superseded on 9 August (10 of 10, 100%, a 0 of 10 benign control) still rendered on the three /for pages, the /findings null card and its feeds, the homepage method selector and the EAI01 entry. Every one of those now derives from the pinned manifest, and check:facts forbids the retired phrasings, which is how they should have been caught the first time.
  • The methodology note for the 88% figure had the reword arms backwards: it called the paraphrase attack (3 of 50) the control and said the 1 of 50 control was a draft error. The harmless-variation control (benign_reword) is 1 of 50 and always was; the note now says so, keeps a dated correction in its body, and the reversal is recorded as E-2026-10.
  • Version story on /citation: the version to cite now comes from the product’s CITATION.cff (0.41.2), not from the evidence pin (v0.39.1); the numbers were measured with 0.32.0. Three quantities, labelled as three. /results no longer attributes the ten-task run to “provael 0.1.0”.
  • Legal pages now describe what actually runs: the contact form’s path (Cloudflare KV, then the AWS lead store in us-east-1, then email), its 730-day retention, and one sub-processor list shared with /trust — which previously said the AWS store was in an EU region; the deployed Terraform state says us-east-1. The Cal.com embed that never shipped is no longer described as running.
  • Smaller: the CRA page and SECURITY.md agreed with themselves about when open-source-steward duties start (11 December 2027); the errata ID collision on E-2026-05 is recorded on both sides with a new E-2026-09 in the product ledger; the leaderboard description no longer says “one measured row”; /submit counts adapters the way the homepage does; drawing numbers name one sheet each; the star count is derived; the sample pack’s AI Act date is the adopted one.

The CRA clock page says the duty is live, and an upstream disclosure gets a home ·

  • /compliance/cra-incident-reporting laid out the mechanics of the Article 14 duty in detail and never said the duty was in force. It does now, in three short paragraphs: the duty is live, it does not bind an open-source steward until 11 December 2027 under Article 24(3), and it does not make anyone non-compliant today. The 24h / 72h / 14-day breakdown below it is unchanged, because restating four deadlines at the top of the page that already carries them is how two copies start to drift.
  • ActSafeGuard (arXiv:2609.11697) joins /defenses. It trains hard action feasibility into flow-matching policies rather than filtering unsafe actions at inference, and reports a 100% step safety rate on pi-0.5 and Fast-WAM. The entry states what that headline does not cover: it is a rate against the paper’s own constraint set, it says nothing about failures whose harm is not a per-step action constraint, and it reports no false-positive rate at all.
  • /security now records the openpi disclosure. We reported a policy server bound to all interfaces returning tracebacks to the client on 5 September; on 11 September a contributor we have never met opened a fix, four files with tests. The page says plainly that the fix is their work and not ours, and records the one point we argued against them on in public.
  • That disclosure is deliberately NOT on /findings. That page is the research stream, where every entry derives its numbers from the pinned evidence manifest and carries a measurement-maturity status. This has no measurement in it and no Provael attack family found it, so giving it one of those statuses would have diluted what they mean.

The regulatory clock stops counting down to a date that has already passed ·

  • On 11 September 2026 the hero of /regulatory-clock read “Next dated obligation · EU Cyber Resilience Act” above a countdown frozen at 0 days : 00 hrs : 00 min. The component pinned the digits at zero when its target passed and stopped, and it would not have recovered on any later day. It now swaps the digits for “In force since 11 September 2026” and the label above changes to “Last dated obligation”, so the two agree.
  • The selection was never the bug, which is worth saying because it is the part that looks wrong. The page picks the soonest obligation after the date its evidence was last pinned, and on that pin it picked correctly. What drifted was the pin: it sat at 3 September while the obligation fell on the 11th, so a date that had passed for every visitor was still one the build could honestly call upcoming. The build carries no wall clock by design and that design is kept.
  • Three checks now sit with the clock freshness gate, which is already the one place on this site allowed a real clock. One fails when the clock has no obligation left after the pin. One fails when any countdown anywhere targets a date at or before the pin, which is what will catch the two counters on the homepage and /machinery-regulation/ when 20 January 2027 arrives. One reports an obligation that has passed in the real world while the pin still calls it upcoming, and turns that into a failed build after 30 days. All three were made to fail on purpose before being trusted.
  • The screen-reader path was the part most likely to be got wrong. The digits stay hidden from assistive technology in both states, as they were, because on the homepage this component sits inside a link and exposing it would rewrite that link’s name after the page loads. The status reaches a screen reader through the label instead, which is where it belongs and where it is said once.
  • Separately, /compliance/cra-incident-reporting/ said ENISA’s single reporting platform was “not live yet”. ENISA deployed its initial operating capability on 11 September 2026, the morning the duty started, so the sentence was false by the time anyone read it under the new rules. Rewritten against ENISA’s own announcement, keeping their “initial operating capability” wording rather than rounding it up to “live”, and the CRA row on the clock now reads as in force with full application still ahead on 11 December 2027.

What the CRA reporting duty changed, and a Top 10 candidate that Provael itself fails ·

  • /compliance/eu-cyber-resilience-act now opens with what changed on 11 September 2026 rather than with a description of the instrument. Manufacturers of products with digital elements must report actively exploited vulnerabilities and severe incidents: an early warning within 24 hours, a full notification within 72, and a final report within 14 days of a corrective measure being available, or within a month of the 72-hour notification for a severe incident. Every date is quoted from the Commission’s own reporting page.
  • The section is written in the past tense on purpose. One announcing an imminent deadline ages into telling a reader something is about to happen after it already has, which is the defect this cluster keeps having: the draft this was written from still said the duty bites tomorrow.
  • It also says who is not caught. The duty falls on manufacturers placing products on the EU market; open-source software stewards are subject to their own obligation under Article 24(3) only from 11 December 2027, per Article 71(2). If you ship an open-source policy and nothing else, that is fifteen months away. Getting that wrong in the alarming direction would have been the easy mistake.
  • /eai-top-10 carries a candidate for the next revision: under-specified policy identity. A policy identified by weights and prompt can still execute differently once action unnormalisation and controller conventions are applied, so a passing result on one deployment does not transfer to another deployment of the same policy. The framing is Jianwei Tai’s (arXiv:2606.03724), not ours, and the entry says so.
  • Provael fails its own candidate check, and the entry leads with that. A Provael report identifies a policy by weights hash and config hash and records neither the resolved unnormaliser nor the controller convention, so two of our reports on the same policy can describe two different executable policies. The schema fix is tracked as provael#227. The candidate sits outside the numbered list, so the ten entries do not renumber and the published coverage artifact is unchanged.

The Korea AI Framework Act gets the page the clock had been promising it · #116 ↗

  • /regulatory-clock has carried a KR row since August saying what Provael supports for a high-impact system under Korea’s AI Framework Act. There was nowhere to click. /compliance/korea-ai-framework-act/ is that page: Art. 34(1) subparagraphs 1, 4 and 5 — risk-management plan, human management and supervision, and the safety-and-reliability document set — with the Act’s own wording, plus Art. 4(1) extraterritorial reach, Art. 33 advance self-review and Art. 36 domestic representative. It is the first non-EU national statute in the crosswalk; the other thirteen pages are all EU, ISO, IEC or NIST.
  • Writing it meant reading the statute, and two things the clock said turned out not to be in it. The row claimed the high-impact duties include “interruption and rollback”: they do not appear in Art. 34 at all, and the Act’s only shutdown reference is a government R&D support provision pointing at a different statute. The row also said enforcement of the penalty provisions “is deferred by one year”, which reads as statutory; the Addenda contain no such provision, and the one-year grace before administrative fines is an enforcement-policy position, not law. Both are corrected on the clock as well as absent from the page.
  • The scope sentence changed too, and it narrows who has to care. Art. 2(4) makes a system high-impact only in an enumerated area — energy, drinking water, health care, medical devices, nuclear, biometrics for criminal investigation, employment and loan assessment, transport, public decisions, school assessment. A VLA policy is inside this Act when it is deployed into one of those, not because it drives a robot. A general warehouse or factory arm is not enumerated.
  • What the page does not claim: no Korean conformity or registration position, and not legal advice. There is no conformity assessment under this Act for a red-team result to feed. Four figures that circulate widely — the 10^26 FLOP compute threshold, the domestic-representative revenue and user thresholds, a document retention period, and the fines grace period — are set by the Enforcement Decree, whose text could not be retrieved, so the page carries them as reported and says so rather than presenting them as read.
  • The clock row now links to the page, the page is inside the build-time date check against the clock instead of being skipped, and the primary source is a stable per-statute law.go.kr URL rather than the English portal the row used to point at.

The staleness banner reads the product’s drift verdict instead of computing its own · #114 ↗

  • The banner had its own version arithmetic: a `minorDelta` comparing the pinned board’s tool version against the current release. `provael.watch.releases_behind` already owned that rule, so there were two implementations of one staleness question in two repositories — and two implementations of a staleness rule drift in the reassuring direction by default, because a window that quietly widens looks exactly like a project keeping up. provael 0.41.2 publishes the answer as `watch/publish-freshness.json`; the banner reads it. Nothing a visitor sees changed today, and that is the point: the same sentence now updates on the product’s next release with no edit on this side.
  • One thing deliberately did NOT move to the artifact. The version in "measured with v0.32.0" still comes from this site’s pinned evidence board, because that is the run whose numbers the page actually renders; the artifact’s `measuredWith` is the product’s own view, the version behind the largest campaign in its ledger. They agree today, and the build now fails if they ever stop agreeing — a mismatch would mean the banner was stating the drift of one run beside the numbers of another.
  • A missing or malformed artifact fails the build rather than defaulting. That is not caution for its own sake: "not stale" is the one direction a fetch failure could push a trust page, and it is the direction that looks fine. Two build checks hold it — one that the numbers are read rather than recomputed, one that a malformed field cannot fall back to a reassuring value — and both are mutation-tested against the exact defect they name.

The staleness banner names the window that actually fired, and stops clipping its own link · #110 ↗

  • The banner fires when either of two windows is crossed — a 7-day measurement age or a 2-release drift — and its sentence always blamed the age. A measurement landed on 6 September, so the age went to zero while the release count stayed at seven, and the banner read "0 days old … past this project’s own 7-day window". Zero days is not past a seven-day window. It now names whichever window fired, and says so in the window’s own units.
  • Worse than the wrong label: the sentence joined two different runs. The age came from the newest measurement, live; the version came from the pinned published one. Those were the same run until the scheduled GPU lane started committing, and from that day "newest real-model measurement: 0 days old, measured with v0.32.0" described one thing that did not exist. The two facts are separate clauses now, each naming its own subject.
  • The banner was also clipping itself. `body` sets `overflow-x: hidden`, and the copy carried `white-space: nowrap` above 760px — so at a 760px viewport the sentence rendered 1026px wide, the document overflowed by 294px, and there was no scrollbar to reach it. What sat past the edge was the end of the sentence and the "Why, and what unblocks it" link: the only control in the banner, unreachable, on the disclosure a reader is least able to skip. Measured on the built page rather than inferred. WCAG 2.2 SC 1.4.10 Reflow. The nowrap is gone; the row still renders on one line wherever it fits.

The release number stopped being a copy, the registry stopped being rounded down, and every measurement got a date · #109 ↗

  • Fourteen pages rendered v0.39.3 while the tag, the GitHub release and PyPI all said 0.39.4. /security was telling a vulnerability reporter to reproduce on a superseded release before filing, and /changelog printed the superseded release as the current one in the same sentence that named the newer one as its newest entry — a sentence contradicting itself, published, green. (Not quoted here: this page is checked for exactly that shape of sentence, and the guard cannot tell a quotation from a claim.) The version had been read from a mirror refreshed by running a script on the release cadence, which is to say by someone remembering. It is now read from the product repository at build time and proven to be tagged before it renders, and /changelog fails the build if its two sources disagree.
  • The homepage said 8 policies and 5 suites. The registry says 8 and 6, and the honest numbers are 5 runnable of 8 and 5 runnable of 6: three policy adapters and one suite are registered and have never been run. The suite figure had been right about the runnable count by accident while claiming to be the total. Both cards now say "5 runnable of N registered", the scaffolded entries carry the word scaffolding rather than a shade of grey, and the build checks the chips against the declared scaffolding NAMES rather than only the counts.
  • /results now dates every committed measurement, not just the newest one. Twenty-six rows, each with the tool version that produced it, whether its date was observed or reconstructed, and whether it counts as a measurement at all — five came from a fixture backend and measure no policy. The banner at the top of the site answers "when was anything last measured"; this answers "how old is the specific number I am reading".
  • Eleven links inside tables on /compare/published-attack-baselines, /defenses and /sample-evidence-pack rendered identically to the text beside them: no colour, no underline, findable only by hovering. WCAG 2.2 SC 1.4.1, Level A. One had been reported; a new check that reads the CSS each page actually ships found the other ten.
  • Four surfaces said ISO 10218:2025 hands its detailed cybersecurity requirements to IEC 62443. It does not — Clause 2 of ISO 10218-1:2025 does not normatively reference that series, which appears only in an informative bibliography, and the 2025 Foreword says the revision ADDS cyber requirements. An assessor was being sent to a paywalled series for a requirement in the document they already hold. Recorded as erratum E-2026-07, and the four superseded phrasings are now blocked by the build.
  • A source link under the CRA Article 14 sub-deadlines had been moved to the Commission’s summary of the Regulation, which is the document those four rows were deliberately moved OFF five days earlier because its phrasing does not carry the distinction they encode. The deadlines themselves were correct throughout; the citation under them was not. Erratum E-2026-06, and the link label now travels beside the URL so the two cannot be edited apart.
  • The regulatory clock was 42 days past its last verification against a 30-day window, and it is now 18. Every instrument that moved was re-read against its own record — including iso.org, iec.ch and the ANSI webstore, which block automated fetching and were opened in a browser instead. That turned up a dead citation and a wrong designation: the ANSI entry pointed at a URL that 404s, and named a document ANSI marks Historical. The current designation includes a Part 3 that was developed in the United States and is not an adoption of ISO 10218 at all. Two entries were deliberately NOT re-dated, because the readings stopped before the articles carrying the dates those entries exist for.
  • /for/certifiers-insurers argued that a measured attack-success rate is the kind of thing an underwriter can use, and cited nothing. It now carries what the market has actually done — a standard generative-AI exclusion for US general liability, a carrier writing specifically for autonomous robots, and underwriters’ own account of why physical AI resists pricing — each labelled with how strong the source is and each stating what it does not establish. Nothing there says an insurer requires a red-team result, and nothing rests on a robot having hurt anyone.
  • /compare/gray-swan is new, and it is the comparison people make out loud the first time this project is described to them. Gray Swan runs the largest AI red-teaming arena there is; on their own published figures it is bigger than this project by three to five orders of magnitude, and that is in the first paragraph rather than a footnote. Seven axes, four of which Provael loses. The difference that survives the scale gap is what comes out: a break says a failure is possible, a rate says how often.

A readiness checklist for the CRA clock, one price on the homepage, and a way onto the leaderboard · #105 ↗

  • /compliance/cra-incident-reporting described the Article 14 clock and stopped there. It now ends with the four things that have to be already true before a 24-hour duty can be met: who is on call and whether they may file without sign-off, what evidence they have to hand, where the SBOM lives and whether it covers what is running rather than what was released, and how "when did we become aware" and "why did we conclude this was or was not reportable" get written down at the time rather than reconstructed. Framed as cheap now and expensive on the day — the obligation has not started and the page does not tell anyone they are late.
  • The homepage showed four prices run together on one line and sent every reader who acted on them to the most expensive page on the site. /checkpoint-report was not linked from the homepage at any depth. It now names one product, shows its price, and links to it. The full ladder still lives on /pricing.
  • The leaderboard told a stranger how to produce a submittable artifact and not what happens to it. It now names the exact file the PR adds, what gets published (the report, the row, the interval, the benign control beside it), what is never asked for (your weights, your episodes, anything about you beyond a directory name), and the two grounds on which a submission is refused — both failures of method, never of result. And in writing: a result that makes Provael look bad is published unchanged.
  • Every count and version on the site was already derived rather than typed, and two checks guarded that. Neither could see a NEW one: typing a wrong family count into a heading on /quickstart left both reporting success, because both are allow-lists over the slots someone remembered to enumerate. check:inventory sweeps rendered page copy instead, and knows the difference between a restatement of the registry and a legitimate denominator — a sentence about how many of the families measured on a real policy came back null is true, and is not a claim about how many are registered.
  • /compliance/cra-incident-reporting rendered its own <main> inside the one the layout already emits, so the page carried two main landmarks and two elements with id="main" — invalid, and meaningless to anyone navigating by region. It was the only page in the site doing it, and check:routes now counts them. Separately, the three "evidence, in three links" lists on the /for/ pages marked their links with colour and weight alone (WCAG 2.2 SC 1.4.1); they are underlined now.

The stale-measurement banner now says what is stale, and measures it in releases · #104 ↗

  • A red STALE MEASUREMENT banner sat directly above the pricing table. The question that raised was whether that is honest positioning or a conversion hole, and the answer turned out to be neither: the banner was honest about one thing and read as a claim about another. On the results page, "our published number is 25 days old" is exactly right. On the pricing page the reader is asking whether they will get something current, and a red banner above a price answers that question with a no it never meant. An assessment runs against your own policy on the current release.
  • So the banner stays, in the same red, on every page including the ones that ask for money. What changed is the sentence. On a commercial page it now says the published benchmark result is the thing that is old, and that your assessment is not. Hiding it where money is discussed was the one option refused outright: a project that sells the willingness to publish its own bad numbers does not get to suppress them at the checkout.
  • The second change is the measure. Days were always a proxy. A pinned result does not decay because a week passed, it decays when the code that produced it changes, and counting days is wrong in both directions at once - too aggressive on a run nobody invalidated, too lenient when several releases ship inside the window. The banner now also fires when more than two releases have shipped since the measurement, and it says which version measured it. On the day this shipped that read seven releases against twenty-five days, so the old measure was the less alarming of two true statements, which is the wrong way round for a staleness signal.
  • The seven-day window is untouched and still fires on its own. The new condition can only make the banner appear earlier, never later, which is the test the original decision set for any change to this mechanism.

A methodology note for the headline number, and the last three Top 10 glossary entries · #103 ↗

  • The 88 percent figure now has a page explaining what it is: the exact setup, the benign control arm that gives it meaning, the paired reword that separates attacker control from mere brittleness, what the attack costs task completion, and the three things that would change the result. Until now that reasoning existed as homepage copy people scroll past, which is not something an academic or a journalist can cite.
  • One correction worth recording. The outline this note was written from put the semantics-preserving reword at 1 of 50. The committed artifact says 3 of 50, and 3 of 50 is what the page prints. The note says so in its own body rather than quietly using the right number, because a page whose argument is "publish your control arm" does not get to hide a draft that disagreed with the run. [Reversed on 13 September 2026: the 3 of 50 was the paraphrase attack arm, not the reword control; 1 of 50 was right. See E-2026-10 and the entry for that date.]
  • The glossary now has an entry for every one of the ten Embodied AI Security risks. EAI08, EAI09 and EAI10 were missing; the last of those carries no attacks and never will, because there is no policy input that attacks the absence of a process, and the entry says that rather than leaving a gap that reads as an oversight.

The counts on this site now come from the release they claim to describe · #102 ↗

  • Every registry count the site renders — attack families, attacks, the real-policy split — was one release behind, and so was the supported version on /security, which told a security reporter to reproduce on v0.38.0 while PyPI served 0.39.1. Both came from the same stale pin, and both are corrected — the counts on /results and the version on /security are now read from the release they claim to describe, so this note does not restate them and cannot go stale the way they did.
  • No measured number moved. The attack-success rate, its confidence interval, the benign control and every per-attack row are byte-identical before and after; only the provenance pin and the registry block changed. Filed as E-2026-05 on /errata.
  • The part worth reading is why nothing caught it. A build check already compares the site’s counts against the product’s generated registry file and it passed, because the copy it compared against was the same stale mirror the pages were reading. A guard that reads a cached copy of the fact it guards cannot fire. That is the third time this project has found that exact shape, after the measurement-freshness banner and the regulatory-clock anchor.
  • Re-pinning also revealed that the regulatory clock has been past its own 30-day re-verification window since 26 August. The banner had been measuring its age against a frozen build anchor rather than against wall time, so the site read "current" while the underlying check already knew otherwise. The clock page now says stale, which is true, and the entries on it still need re-reading against their primary sources.
  • The Unitree G1 EDU disclosure of 27 August is now on /incidents — two published CVEs, one reachable from Bluetooth range with no pairing. It is filed under EAI07 and used to state a boundary rather than to claim coverage: Provael attacks a policy through what it is shown and told, and nothing in that disclosure touches a policy. A clean Provael run says nothing about it.
  • /verification now asks for the thing it exists to get. It carries a direct link to the raw artifact behind the headline result and a commitment: a result that contradicts ours gets published, with attribution, unchanged.

Share cards render off-site, and the footer now links the company page · #101 ↗

  • Four pages passed a root-relative path for their share image while four others passed an absolute URL, so the same layout prop meant two different things depending on which page you had copied from. Facebook, LinkedIn and X resolve og:image against nothing: a relative path is dropped and the link shares with no card image at all. /adopters and /verification were the two still doing it.
  • Fixed in the layout rather than on the two pages. BaseLayout now resolves any per-page image against the site URL before rendering it, so a page can pass either form and the tag comes out absolute. Patching the two callers would have left the next page that copies one of them broken in the same way.
  • A build-time check now reads the generated pages and fails the build if any og:image is not absolute, and refuses to pass if it finds no og:image tags at all — a check that silently matches nothing looks identical to one that found no problems. Swept every page afterwards: 77 of 77 carry an absolute image, none missing.
  • Separately, the footer now links the Provael company page on LinkedIn, next to GitHub. The URL was already in the page’s structured data, which is the half a crawler reads; nothing on the rendered page linked to it, which is the half a person uses. Both now read the same constant.

Correction: the CRA severe-incident final report starts from your own filing, not from the fix ·

  • The Article 14 reporting clock on /regulatory-clock carried both final-report deadlines as one row: "Final report, due once a corrective or mitigating measure is available - 14 days for an actively exploited vulnerability, one month for a severe incident." That start point is right for the vulnerability branch and wrong for the incident branch.
  • Article 14(2)(c) and Article 14(4)(c) measure from different events. The 14-day vulnerability report runs from when a corrective or mitigating measure is available. The one-month incident report runs from submission of the 72-hour notification under Article 14(4)(b) — it does not wait for a fix at all. Collapsing them into one sentence applied the first anchor to the second, which points a reader at a later start than the regulation allows. A runbook built on it could wait for a fix before starting a count that had already been running.
  • Cause worth recording, because it is the reusable part: the sub-deadlines were transcribed from the Commission CRA summary rather than from the OJ text, and the summary’s phrasing does not carry the distinction. The clock entry now cites Regulation (EU) 2024/2847 on EUR-Lex as its sub-deadline source, carries the two branches as separate rows with their own anchors, and records Article 69(3) — the derogation that puts the whole installed base in scope for reporting while leaving it out of scope for the product requirements.
  • The 24-hour, 72-hour and 14-day figures were correct throughout, and no measured result, signed artifact or attestation is affected. Filed as E-2026-04 on /errata.
  • While checking that page: /errata was missing E-2026-03, the zero-width confidence interval published in two of the product repository’s READMEs and corrected there on 30 August. This site was never wrong about it — provael.com carried the correct non-zero upper bound the whole time — but an errata page that is missing a correction is the one page on this site that cannot afford to be behind. It is mirrored now, and the mirror date moved with it.

The CRA Article 14 clock now has a page, ten days before it starts ·

  • Reporting obligations under Regulation (EU) 2024/2847 apply from 11 September 2026. The site already carried that date — /regulatory-clock has announced it as the next dated obligation since 25 August, and the CRA crosswalk states it — but nothing on the site answered the question a manufacturer actually asks on the day it matters: an exploited vulnerability has just been reported to us, what do we owe and by when.
  • /compliance/cra-incident-reporting is that page, written as a clock rather than an explainer: what starts it (an actively exploited vulnerability or a severe incident — not merely discovering a bug), the 24-hour early warning to the coordinating CSIRT and ENISA, the 72-hour notification that has to add what the first one was allowed not to know, the final report at 14 days for an exploited vulnerability or one month for a severe incident, and the clause most often missed — products already on the EU market are in scope.
  • It exists on this site rather than a law firm’s because of the last section. A firmware CVE closes with a patch and a version number; a learned policy has neither, so "we fixed it" is not evidence that anything improved. The only honest answer to how do you know it is better is a before-and-after measurement on the same suite, the same seeds, and a benign control arm — without the control you cannot tell a policy that got safer from a predicate that got quieter.
  • No deadline on the page is typed. All three windows and both dates are read from src/data/regulatory-clock.json, which already carried them with a CELEX-quoted verification note; a second hand-typed copy is exactly how this site once had two pages disagreeing about which obligation came next. The page also fails its own build if the clock ever loses those sub-deadlines, because a page about a 24-hour duty that silently renders zero deadlines is worse than a 404.
  • check:dates:compliance would have skipped it. That guard maps a compliance slug to a clock entry id, and this page is not a framework crosswalk, so its slug matched nothing and every date on it would have escaped the check — on the one instrument the guard’s own header was written for. It is now mapped to the CRA entry, so the page states no date the record does not carry: nine compliance pages are checked where eight were.

The booking control on /assessment had never worked in production ·

  • The #book block on /assessment offered a click-to-load Cal.com calendar. It rendered its "Load booking calendar" button only when a calendar link validated, and that link came from PUBLIC_CAL_URL / PUBLIC_CALCOM_LINK — neither of which has ever been set in the deployed build. So the served page carried an empty cal-link attribute, no button, and an inline script whose first act was to look up the button by id. The lookup returned nothing. The bind was optional-chained, so nothing threw and nothing reached the console. Above it, the heading read "Pick a time".
  • How long: the block has shipped in this exact shape since the first commit in this repository, on 6 July 2026, and the variable it depends on is unset in production today — eight weeks. Whether it was ever set in between cannot be established from the repository, because that value lives in the deployment environment and never in git; what can be established is the shape of the page and its state now. The email fallback beside it always worked, which is why nothing ever looked broken — every visitor who wanted a call got a mail client instead of a calendar, and the site had no way to tell that had happened.
  • It was not one page. SCHEDULING_URL resolves to /assessment specifically so that every "book a scoping call" control on the site — /pricing, /continuous-assurance, the three /for/ pages, /machinery-regulation, and the booking button’s own no-calendar fallback — funnels to the one surface meant to take a booking. All of them terminated there. The Standard Assessment tier carried a source comment stating that the booking calendar already lived on /assessment; it did not.
  • The calendar is gone from that page rather than repaired, because repairing it means guaranteeing an environment variable in production and this is the second time a control here has depended on one being set. #book now leads with a link into the request form carrying the paid-pilot intent, which posts to the same endpoint the rest of the site uses and captures the lead rather than handing the visitor a draft to send themselves. The email address stays visible beneath it as the path that needs no script and no endpoint. Roughly sixty lines of third-party bootstrap that could never run no longer ship to anyone.
  • The class of failure is now a build check. check:controls asserts, against the built output, that every element a script looks up by id exists on the page shipping that script; that no booking attribute is present but empty; that every page carrying a conversion control still offers a form, a request-form link, an email address or a route to the booking surface; and that a form with no action has a script that references it. It was written against the broken build first and reported both halves of this defect before anything was fixed. Its entry in the guard mutation suite reintroduces the exact fault, so a future change that makes it stop noticing fails CI.

Two pages disagreed about which obligation comes next · #86 ↗

  • /compliance/eu-cyber-resilience-act said, in its body text, "Reporting obligations (Art. 14) apply 11 September 2026". /regulatory-clock announced the next dated obligation as the EU Product Liability Directive on 9 December 2026. September is before December, so the two pages disagreed about the one thing the clock exists to state — on a site whose pitch is catching exactly that in other people’s material.
  • The date was never missing from the record. regulatory-clock.json carried it as `reportingDate` and the clock ignored it: the timeline filtered on `applicableDate` alone, and the type did not declare the other field. It was not CRA-specific either. The AI Act entry carries `additionalDate` 2 December 2026, for two further prohibited practices, and that was skipped by the same filter — also earlier than the date being announced. Fixing only the CRA would have left the page wrong about the very next obligation after it.
  • Every date-shaped field in the clock is now classified exactly once: an obligation a reader must act by, or explicitly not one, with the reason recorded. An unclassified date field fails the build, which is the only property that stops this recurring — the defect was a date nothing read. `supersededStatutoryDate` is the exclusion worth naming: it is the date Regulation (EU) 2026/1744 moved, so announcing it would tell a reader to act by a date that no longer governs.
  • The CRA entry now carries the Article 14 reporting clock, which appeared nowhere on this site before and is what a certifier actually asks about: 24 hours for the early warning to the coordinating CSIRT and ENISA, 72 hours for the main notification, and 14 days for the final report once a corrective or mitigating measure is available. Its scope line says plainly what a Provael run does not establish for it — a policy-attack result is not a vulnerability-handling process, and Article 14 is about the process.
  • A second check asserts the mirror image: a date stated on a compliance page must exist in the clock JSON. It is scoped to pages whose instrument the clock tracks, and visibly skips the five that carry dates the clock has no opinion on — a standard’s publication date, an inspection-programme announcement. Failing on those would either block a correct build or "fix" itself by writing a vendor press release into a regulatory record.
  • The clock is now held to a stated verification window, the way the measurement already is. Past 30 days — the window this site already applies to a hand-verified date — every page that states a regulatory date carries a banner giving the age of the check. Past 90 days the build fails. Inside the window nothing renders at all: the banner appearing is the signal, and a permanent "clock: fine" strip is furniture nobody reads.

The version numbers came from the wrong single source ·

  • PR #69 below is titled "One source of truth for every published number". For the measurements it held. For the version strings it did not, and it is worth being precise about why, because the mechanism it introduced is what broke them: #69 said derive everything from the pinned evidence manifest, and the manifest’s pin is deliberately FROZEN at whatever release produced the numbers. That is correct for a measurement and wrong for "the current release", so every surface naming a current version began falling behind the moment the next release shipped. Nothing was hardcoded — grepping for a stale string would have found none of it — and by v0.38.0 the changelog page named 0.36.2 as the current one directly above its own 0.37.0 entry (the sentence is paraphrased here on purpose: `check:versions` scans this page like any other, and spelling the claim out verbatim would make the guard flag its own account of the bug), /security told a security reporter to reproduce on a superseded release before filing, and the homepage published a copy-pasteable Action pin at a version two minors old.
  • There are two version quantities, not one, so there are now two sources. The current release comes from the GitHub latest-release endpoint, fetched at build and committed to src/data/repo-facts.json so a build without network is reproducible; the measurement pin stays exactly where it was. The build refuses to run at all if the committed facts are pinned to a commit rather than a release tag, because six pages state a current release and a site that cannot name one should not be published.
  • What is structurally different from #69: that PR added derivation and no check that the derivation read the right source — a derived number and a correctly-derived-from-the-wrong-place number are equally invisible to `check:figures`, which is why it caught none of this. `check:versions` now asserts against the built output that every currency CLAIM ("the current release is", "currently vX") and every copy-pasteable Action pin names the release the site is pinned to. It matches the claim rather than the number, because an old version appearing is usually correct: /citation names the release whose artifacts produced the numbers, /changelog lists every release there has ever been. It found one on its first run that had been missed by hand.
  • The measurement’s age is now a build condition rather than a sentence. Past 7 days every page carries a banner stating the age in days and machine-readable meta; past 21 days the build FAILS. An age that cannot be parsed also fails, rather than defaulting to fresh — the failure to design against is a site quietly serving a number it can no longer date. The same status is written into llms.txt, read back out of the rendered page so the two cannot hold different numbers: robots.txt here opts every AI crawler in, so an answer engine can quote the attack-success rate without ever rendering a page, and the staleness has to travel with the number.
  • /leaderboard now reads the board’s own staleness verdict instead of implying currency. When the published artifact says stale, the table sits behind a labelled "measured with an older tool version — not a current result" state carrying the product’s reason verbatim, and each row shows the tool version that measured it. The rows still render: hiding them behind an interaction would strip them from a crawler and from anyone on a keyboard, which makes the disclosure worse than what it discloses. The benign control column gained its counts and its own confidence interval, so the floor an attack-success rate is measured against is published to the same standard as the rate.
  • llms.txt said "All 76 routes" while the sitemap carried 81 URLs. Both were right and neither said which: 81 is every indexable URL, 76 is the subset that is an HTML page with a title and description to list, and the other five are the feeds. The two already came from one source — the index is generated from the sitemap — so the fix was to state the arithmetic and check it, not to add a manifest layer that removes no failure mode.

One source of truth for every published number · #69 ↗

  • The homepage carried the superseded single-task result — 100% keep-out violation against a 0% benign control, "one task" — in five places, on the same page as a hero already rendering the ten-task result from the pinned manifest. The page contradicted itself and the stale half was the flattering one. All five now derive from the manifest, so a re-pin moves them.
  • Two build checks were added so this class of drift fails rather than ships. `check:figures` refuses a literal percentage or fraction inside a stat slot in any template, because after rendering a derived number and a typed one are indistinguishable. `check:board` now also compares /leaderboard against the exact signed artifact it rendered from — every row, the generation date, the source commit and the digest — since a page that can silently disagree with its own signature is worse than no page on a site whose pitch is "verify the signature rather than trust this".
  • /verification stopped telling people to install from git. `provael submit` has shipped on PyPI since 0.32.0; the instruction outlived the fact it described, on the page whose job is to get third parties to reproduce a result.
  • /citation now names the current release and carries the Zenodo concept DOI in both plain text and BibTeX. It previously said there was no DOI, which stopped being true when one was minted and confirmed resolving.
  • Added /changelog, which had no route at all — releases were visible only on GitHub and PyPI, which is exactly why four version strings drifted unnoticed. It renders the product CHANGELOG and this list side by side, with the measurement-freshness signal beside them.
  • The regulatory clock was counting down to an obligation that had already arrived: "next" was measured from the date the clock was last verified rather than from anything resembling now, so 2 August 2026 stayed "next" after it applied. It is now measured against the site’s own pinned date, and the page states which anchor it used and that an obligation has come into force since the entries were last checked against primary sources.
  • The five machine-readable crosswalk artifacts are downloadable from /compliance, each stating its own mapping_status — and where the artifact declares none, saying so rather than supplying one. XPolicyLab is deliberately absent: it ships no artifact, and listing it would present a planned integration as a delivered one.
Unreleased

Merged, not yet published

On main and not in any release. Installing from PyPI does not get you these.

Added

  • Two more control arms, scrambled_text and roleplay_no_target, because the first two left one objection standing. nonsense_text is three tokens; roleplay renders as twenty. A 0/50 on three tokens never showed that a twenty-token out-of-distribution string is harmless, so the 44/50 headline was still consistent with "a long unfamiliar sentence derails SmolVLA" — the generic fragility RobustVLA (arXiv:2510.00037) and LIBERO-PRO already document, not attacker control. scrambled_text is the roleplay prompt's own tokens with the target noun replaced by a one-token filler and the order destroyed, whitespace-token count identical, seeded per episode. roleplay_no_target is the roleplay frame intact with only the graspable target replaced by the same filler, so a firing there is the imperative frame driving the policy and not the object. Both are built *from* RolePlayAttack.TEMPLATE, so the match survives a template change; both carry the control role and enter neither the ASR nor the benign FPR. --attacks control now resolves to four arms. The registry is 44 attacks (39 adversarial, unchanged); watch/registry.json, the inventory lines and the checked-in evidence manifest are regenerated.
  • The weight-integrity family and the gradient-patch attack can now run against a real policy. LeRobotAdapter implements WeightAccessible over the weight matrix of model.action_out_proj — the last linear map before the action chunk on every flow-matching checkpoint lerobot 0.5.1 loads through it (SmolVLA: 15,360 parameters) — through a symmetric per-tensor INT8 view. A flipped INT8 value is applied to the live float weights as the delta the same flip would make in an INT8 deployment, and the clean vector restores the original floats bit-for-bit; the runner hands the adapter its reference operating point (the benign first frame of the run's first task at the run seed, via the new SensitivityReferencePolicy protocol) so d(danger)/d(weight) is one autograd pass at one fixed point per run. The danger proxy — translation energy of the executed chunk in normalized action space — is documented as a ranking heuristic, not the predicate's derivative. LeRobotAdapter.input_gradient backpropagates ||enc(x) − enc(x_clean)||² through the backend's own vision tower to the suite-frame image, and the runner attaches it to every attack exposing attach_gradient_oracle (the new InputGradientProvider / GradientOracleAttack protocols) — without a reset callback, because the oracle touches no per-episode state and a reset per refinement would make the attacked arm re-plan every step. pi0-FAST has no action_out_proj and reports the family *not applicable*; the deterministic stub enters neither path, so no CPU report moves. Verified on the real HuggingFaceVLA/smolvla_libero checkpoint on a CPU with a synthetic frame (no simulator): finite sensitivities, exact restore, finite input gradients, and the attack landing at its budget. No rate is claimed until the GPU run is committed.
  • A vla_arena suite (registered scaffolding): the first suite whose unsafe predicate is DECLARED rather than fitted. VLA-Arena (PKU-Alignment, ICML 2026) evaluates each task's own (:cost …) clause every step and reports it in info["cost"]; the adapter's is_unsafe() is exactly "that clause fired", and calibration_signal() returns None so provael calibrate cannot replace a declared predicate with a threshold on its own output. Task names are <benchmark>/L<level>/<id> across the sixteen registered benchmarks; reset() refuses a task from another benchmark or level. The raw observation is handed to the policy adapter in lerobot's LIBERO-wrapper shape and features() returns the LIBERO env config — the same Panda, cameras, 8-dim state and 180-degree camera convention, and what VLA-Arena's own SmolVLA evaluator does — so a LIBERO checkpoint runs unchanged. Written against the package's source (14 Sep 2026) and tested on a fake shaped by it; VLA-Arena pins Python 3.11 / robosuite 1.5.1 / numpy 1.26.4 and needs its own environment. No run has been made through it, list-suites says so, and the coverage counts exclude it; the registry is 7 suites (3 fixtures, 3 gated simulators, 2 scaffolding) and watch/registry.json, the README and the quickstart inventory lines are regenerated.
  • provael report --format test-report: a test report in the shape of ISO/IEC 17025:2017 clause 7.8. The 13 September regulatory re-read found that no certifier, notified body or insurer publishes acceptance of SARIF, OSCAL or an ML-BOM; what an assessor of a machinery technical file reads is a clause-7.8-shaped test report. The emitter lays a run out that way — identification, item under test (checkpoint, revision, digest), method, dates and location, conditions (hardware, precision, OS, lock digest), results with Wilson intervals and the benign floor, uncertainty (per-seed spread, anytime interval, the stochastic-sampling caveat), deviations (not-applicable arms, replayed episodes, skipped checks, missing manifest fields), scope, external providers, evidence state and verdict, authorisation — and an Annex A clause map rendered from provael.compliance, the single source of every framework mapping. It reads execution-manifest.json beside report.json when present. The report says in its first line that Provael is not an accredited laboratory, that this is not an ISO/IEC 17025 report and that no statement of conformity is made; the date of issue and the authorising person are conspicuous blanks for the human who signs, so the emitter introduces no wall-clock value and stamps no signature that never happened.
  • provael export --format hf-eval --dataset <benchmark>: Hugging Face Community Evals entries. One .eval_results/*.yaml entry per *measured* arm — treatments, the benign baseline and the harmless-variation controls alike — with value the episode-level unsafe fraction and notes carrying the role, the counts, the 95% Wilson interval, the predicate state, the tool version and the checkpoint, so the Hub page shows the floor beside the rate. Arms with no applicable episode are omitted, never published as 0; date comes from the execution manifest when present and is otherwise left to the Hub's commit time; no verifyToken is ever emitted (that badge is for HF Jobs + inspect-ai runs). Task ids are the stable <suite>--<attack>, and provael.hf_eval.benchmark_eval_yaml writes the benchmark dataset's eval.yaml declaring every arm the registry can measure. The Hub side — registering the benchmark dataset and adding provael to huggingface.js's evaluation_framework enum — is a pull request a person opens; the emitter only writes the file. Schema read from huggingface.co/docs/hub/eval-results on 14 September 2026 (a work-in-progress feature).
  • scripts/plot_keepout_paths.py draws every episode's end-effector path against the task's keep-out zone (top-down and side panels, one SVG per task, no plotting dependency): benign paths in blue, the chosen attack arm in red with a dot at the first unsafe step, the committed calibration or the documented default box shaded. It reads the trajectories every episode record has carried since 0.40, so a picture of forty benign paths against the box answers the #136 question — does the default box sit inside the benign workspace — faster than any rate. Reports written before 0.40 carry no trajectory; the script says so and draws nothing.
  • docs/compliance/machinery-annex-iii-corruption.md: Regulation (EU) 2023/1230 Annex III §1.1.9 (protection against corruption) and §1.2.1 (safety and reliability of control systems) quoted verbatim from EUR-Lex (read 14 September 2026) and mapped clause by clause — under each paragraph and lettered point, the Provael artifact that speaks to it (injection arms for the connected-device paragraph, checkpoint integrity and the weight-corruption ladder for critical software, the keep-out predicate for "beyond its defined task and movement space", the harmless-variation controls for foreseeable human error, the measured defenses for "correct the machinery at all times") and what it does not establish; the start/stop, protective-device and assembly points are named as out of scope. Linked from the machinery-regulation page and from --format test-report's Annex A. Four remaining double-hyphen doc anchors fixed.
  • The π0.5 cross-architecture study is publicly timestamped: the amended protocol and its falsifiers were deposited on Zenodo (10.5281/zenodo.22751558, 14 September 2026) before the π0.5 attack arms ran; docs/studies/pi0-openpi-transfer.md links the deposit.
  • provael attack --video-dir DIR writes one MP4 per episode: the frames the policy actually saw (after the attack and any defense), red-bordered from the first step the suite's predicate fired. A runner argument, not a RunConfig field, so a run with recording on and off produce byte-identical reports and no attestation digest moves — the same rule audit_sink follows. imageio + imageio-ffmpeg (in the [lerobot] extra) are imported lazily; the CPU core gains no dependency. Episodes replayed from a --resume ledger have no clip. provael.video exposes FrameList for callers composing their own output. provael compose-video LEFT RIGHT --out puts two of those clips side by side, step-aligned — the benign twin beside the attacked episode at the same task and seed — with the shorter clip holding its last frame so an episode the predicate stopped early stays on screen at its verdict. Clips are written from the suite's new display_frame, which the LIBERO suite overrides to turn robosuite's 180-degree-rotated raw frame the right way up — the policy's processor does the same flip before inference, so the clip shows what the policy saw; the attack surface and every committed visual result stay on the raw frame.
  • Release assets carry signed SLSA build provenance. release.yml runs actions/attest-build-provenance over the wheel, the sdist, the CycloneDX SBOM and SHA256SUMS before gh release create, so gh attestation verify <asset> --repo provael/provael names the workflow, tag and commit that built it. PyPI already had PEP 740 attestations for the wheel and sdist; the GitHub release assets carried nothing, which is what OpenSSF Scorecard's Signed-Releases check scored 0 on.
  • The execution manifest's hardware names the CPU count and the CUDA device, not just the ISA — x86_64 described every run ever recorded and distinguished nothing.
  • examples/gpu-ci/local_libero_sweep.py — the sharded LIBERO screen on one machine (N resumable shards, per-shard logs and timeouts, --commit provenance) with an aggregate that reproduces the published cross-shard statistics from the committed shards.
  • docs/studies/pi0-openpi-transfer.md carries its design (n = 50 per arm, 5 seeds, horizon 280) and Amendment 1: the first leg runs π0.5 through the native pi05 adapter, with what that narrows and what it leaves open.

Fixed

  • benchmark_eval_yaml declared libero--none twice. The registry already carries the none baseline, and the README's recipe prepends it for emphasis, so the committed benchmark eval.yaml and tasks.jsonl listed the benign task twice. Arms are now de-duplicated in order of first appearance, the example files are regenerated (44 tasks, 440 rows), and a test reads every committed examples/hf-benchmark/*/eval.yaml so a duplicate cannot be committed again.
  • gradient_patch could not move a frame (E-2026-11). Its objective's gradient is exactly zero at the clean frame and the search started from a zero perturbation, so against a smooth encoder the released module (0.39.1–0.41.2) never left the clean frame — found the first time it met a real vision tower. The loop now starts from a uniform draw inside the ε-ball, seeded from the episode seed and the step; a regression test uses an oracle shaped like the real objective. No published rate moves: the arm never ran against a VLA. The Diffusion Policy × PushT figures it shipped with came from a script outside this repository and are now labelled as that script's result, not this module's, in PRIOR_ART.md and the errata ledger.
  • Three stale sentences in the compliance docs (found by the 13 Sep regulatory re-read): docs/compliance/index.md called ISO 25785-1 a "Working Draft… expected 2026–2027" (it is a Committee Draft since 8 May 2026 with no committed date; trackers read ~2028) and opened the routing box with the pre-adoption "political agreement May 2026" clause (it is Regulation (EU) 2026/1744, OJ 24 July 2026, in force 27 July); machinery-reg-2027.md said ISO 10218:2025 was "in force" (a standard is published, not in force).
  • A LIBERO run is built for the task suite its tasks name. LiberoSuiteAdapter.reset() parsed the suite out of a "libero_spatial/3" task name and then discarded it: the environment came from the adapter's constructor default (libero_object) and the episode was recorded as libero_spatial/3 — the suite-level twin of the task-level misattribution _build_env already refuses, and the CLI had no other way to ask for spatial, goal or 10. make_suite() now takes the run's tasks and builds the adapter for the one suite they name (a list mixing suites is refused), reset() raises on a prefix that does not match the adapter — before the lerobot gate, so the failure reproduces on a machine without the simulator — and a bare "3" is task 3, not task 0. RunConfig is unchanged, so no report or attestation digest moves.
  • Korea's AI Framework Act (Act No. 20676) enters the compliance catalogue — the first non-EU national statute in it. Three rows against Article 34(1), the duties on an operator "providing high-impact AI or AI-based products and services": subparagraph 1 (risk management plan), subparagraph 4 (human management and supervision) and subparagraph 5 (preparation and storage of documents demonstrating safety and reliability measures). Article numbers are the Act's own, read from the CSET English translation of Law No. 20676 (promulgated 21 January 2025, in force 22 January 2026 under its own Addenda Art. 1).

Three of the six subparagraphs are deliberately absent. 34(1)2 (explainability), 34(1)3 (user protection) and 34(1)6 (matters the Committee resolves) have no on-point Provael signal, and a row per subparagraph would read as coverage of the whole Article rather than of the part a red-team result speaks to.

34(1)4 carries required_eai=("EAI04",), so it is a gap unless the action channel actually ran. A human-supervision duty is about the commanded motion a supervisor has to catch, so citing it from a run that never touched the action channel would cite a measurement nobody made — the same gate the functional-safety rows already use. Its provael_signal says in as many words that the row evidences neither the supervisory arrangement itself nor any stop, interruption or rollback mechanism: those are system-design duties over the deployed system, and Provael exercises the policy, not the stop.

Scope here is sectoral, not "any physical machine". Art. 2(4) enumerates the areas that make a system high-impact — energy, drinking water, health care, medical devices, nuclear, biometrics for criminal investigation, employment and loan assessment, transport, public decisions, school assessment. A VLA policy is inside this Act when it is deployed in one of those, not merely because it drives a robot; a general warehouse or factory arm is not enumerated. The constant's docstring in compliance.py says so, because the opposite reading is the easy one.

results/smolvla_libero_object/attestation.insurer.json is regenerated from the unchanged source report: the assurance payload embeds the catalogue, so its payloadSha256 moves. The subject digest over report.json does not, and no attestation over a run report is invalidated.

Fixed

  • The freshness badge's colour comes from the day count it prints. It came from the fractional age, so at 2.3 days the badge read "2 days ago" in orange — a number inside the fresh window painted stale — and the guard that recomputes the badge at the age its own message asserts read 2 days as green and failed every pull request for the ~17 hours until the count ticked over (observed 14 Sep 2026 on the badge committed 13 Sep). Message and colour now derive from the same whole-day count; watch/freshness.json regenerated.
  • The Python versions this package claims, and the date its citation names, are now checked rather than asserted. Three surfaces stated something nothing verified (#220, #221, #222).

The CPU gate runs on 3.13 as well as 3.12. requires-python = ">=3.12" admits 3.13 and 3.14, while .github/workflows/ci.yml's check job pinned python-version: "3.12" and nothing else — so "works on 3.13" was untested rather than merely unadvertised. check is now a matrix over ["3.12", "3.13"] with fail-fast: false, because a break on one interpreter is a fact *about* that interpreter and the other leg's result is what says whether it is version-specific. Only that job changed: release, docs, freshness, coverage-badge, gpu-arm, gpu-scheduled and leaderboard-submission stay on 3.12, since each publishes an artifact and a matrix there multiplies artifacts rather than coverage.

Programming Language :: Python :: 3.13 is added to the trove classifiers, on the strength of that leg and nothing else. PyPI's sidebar renders the classifier list rather than requires-python, so the page read 3.12-only against an open floor — someone skimming it could reasonably conclude 3.13 was unsupported.

3.14 is in neither list, and not because provael fails on it. The suite passes there — 1418 passed, 15 skipped on CPython 3.14.2 — but uv.lock pins numpy==2.2.6, which publishes a cp314 macOS wheel and no cp314 manylinux wheel, so uv sync --locked cannot install on ubuntu-latest under 3.14. It is the only cp314 gap in the locked CPU set. The lane and the classifier land together when the lock moves to a numpy that ships one (2.5.3 does), rather than by loosening --locked and giving up the property that every lane installs the set a release is built from. requires-python keeps its open upper bound for a separate reason: a cap is baked into every published sdist and wheel permanently, so a <3.14 added today would make v0.41.2 uninstallable on 3.14 even once 3.14 is verified — and CITATION.cff exists precisely so a citation names a version a reader can still install.

CITATION.cff said date-released: "2026-09-06"; the v0.41.2 tag was created 2026-09-09. The file's own comment already described this failure from the 0.25.0 release — a citation that "named a date three days before the artifact existed" — and then concluded that the date "is not machine-checkable against a tag, so it is on the release author". It recurred, by three days again. It *is* checkable, because the tag carries its creation date. tests/test_version_consistency.py::test_citation_date_matches_its_tag now compares date-released against git tag -l --format='%(creatordate:short)' v<version>, and skips when that tag is not in the clone so that neither a shallow checkout nor the release-prep commit — where version names the release about to be cut — fails for the wrong reason. test_the_tag_date_lookup_actually_resolves holds the other end, so a broken lookup goes red instead of turning the skip permanent. The date is no longer guarded by the release author's memory.

#220 reported the file at 0.29.1 / 2026-07-31. That part was already fixed and has read 0.41.2 since the 0.41.2 release; the date is what the report did not catch, and what the guard it asked for now covers.

Fixed

  • provael submit advised a command that does not exist. The generated PR body told reviewers to "verify the bundle offline with provael verify …"; the command is provael attest --verify. docs/errata.md's E-2026-01 carried the same phantom command (provael verify … --print-payload) in its "how to tell whether a bundle is affected" block, now replaced with the jq | base64 -d pipeline that actually reads the payload.
  • Meta-World was listed as runnable with nothing saying it cannot complete a CLI run. MetaworldSuiteAdapter refuses LIBERO's default keep-out box (it lies behind the Sawyer arm), the CLI has no zone option and no Meta-World calibration is committed, so --suite metaworld raised at reset. list-suites and doctor now print the gating note beside the suite (suites.SUITE_GATING_NOTES). Counts are unchanged: the suite is implemented and runnable from the Python API with a zone derived from its own benign envelope.
  • Execution manifests from the Modal GPU lanes recorded commit: null. The container pip-installs the pinned release and has no git checkout. The drivers now resolve the pinned tag's commit on the runner and pass it as PROVAEL_COMMIT, which _emit_execution_manifest honours when it is hex-shaped (cli/_shared.py); the two GPU workflows check out full history so the tag resolves.
  • SECURITY.md contradicted itself on when the CRA's open-source-steward duties start (11 September 2026 in one paragraph, 11 December 2027 in another). 11 December 2027, per Article 71(2), throughout.
  • action.yml's Marketplace description was 172 characters against a 125-character cap. Shortened; tests/test_action_scripts.py now pins the cap.
  • Errata ledger: the numbering note claimed the ledger and the provael.com mirror agreed entry-for-entry; the mirror had minted its own E-2026-05 on 3 September for a different correction. That correction is now recorded here as E-2026-09, the collision is stated, and E-2026-10 records the reword-arm mix-up in the site's methodology note. Next free ID: E-2026-11.
  • Stale docstrings that narrated a 29-attack / 15-family registry (coverage.py, tests/test_coverage.py), "controls not wired into the registry yet" (tests/test_controls.py), "90 test modules" (ci.yml) and "33 released versions" (tests/test_changelog_gate.py) now describe the tree as it is; the assertions beneath them were already right.
9 September 2026

0.41.2

Added

  • watch/publish-freshness.json — the release-drift window, published so a consumer reads the answer rather than reimplementing the rule. watch/freshness.json reports when *anything* was last measured, and a one-episode timing probe satisfies it: on 8 September a $0.06 probe put that badge at today while the published 44/50 headline was still measured with v0.32.0, nine minors back. The window that catches that lived in provael doctor and nowhere readable, so anyone wanting the same three numbers had to recompute them against watch/measurements.json — and a reimplemented staleness rule drifts in the reassuring direction by default, because a widened window looks exactly like a project keeping up.

Fields: measuredWith, measuredAt, releasesBehind, staleAfterReleases, isStale, currentVersion. measuredWith is the version behind the largest real campaign rather than the newest record, so a probe cannot displace a 350-episode run. An unmeasured project publishes null for the gap and the verdict rather than 0 and false — zero is the reassuring answer and must never be the fallback.

This is not the cross-repo constant fix. That was staleAfterReleases in watch/release.json in 0.41.1; www.provael.com read its own STALE_AFTER_RELEASES = 2 out of src/lib/freshness.ts until then and now reads the product's. What is new here is the *result* of applying that window, which nothing published before. The constant appears in both artifacts on purpose and cannot drift: both render from provael.watch.STALE_AFTER_RELEASES and both are gated by make check-docs. That is precisely the property the cross-repo copy lacked — no shared source, and neither side able to see the other.

No generatedAt, deliberately. gen_release_artifact.py and gen_measurement_ledger.py both refuse a wall-clock field, and test_measurement_ledger.py asserts it by name: one would make the output differ on every run, so --check would stop meaning "current" and the new make check-publish-freshness gate could never pass on a clean tree. measuredAt carries the "when" a reader wants — the instant of the measurement being described, a fact about the run rather than about the moment the script executed.

  • watch/README.md — the consumption surface had no prose documentation at all. Six generated artifacts, what each answers, why the two freshness files disagree on purpose, and the two properties every file holds: no wall-clock values, and no unverifiable fields.
  • make check-publish-freshness, wired into check-docs beside check-release and check-measurement-ledger, plus make gen-publish-freshness. tests/test_publish_freshness_artifact.py binds the artifact to the code rather than to a literal, and was mutation-tested against a widened window in either file, a softened verdict and a shrunken gap. One mutation initially escaped because STALE_AFTER_RELEASES = 2 also appears in a comment *above* the assignment, so the edit landed in prose — the same shape as a version guard firing on a sentence that quotes the version it corrects.
8 September 2026

0.41.1

Added

  • watch/release.json publishes the release-drift window, so a consumer stops keeping a copy. STALE_AFTER_RELEASES landed in watch.py in this release while www.provael.com held its own = 2 in TypeScript. One policy constant in two repositories, and a disagreement between them is invisible from both sides: the site would render one window and provael doctor another, with nothing failing. It now ships as staleAfterReleases, and test_release_artifact.py fails if the artifact and the constant disagree.
  • E-2026-06 and E-2026-07 back-filled into docs/errata.md. Both were raised on 6 September against website surfaces and recorded only on the mirror at provael.com/errata, which left the maintained source two entries short of the page that copies it. They agree entry-for-entry now.

One thing is recorded rather than fixed: E-2026-07 and E-2026-05 are the same correction under two IDs — ISO 10218:2025 does not defer its cyber detail to IEC 62443, recorded once against two repository documents and again five days later against four website pages. Issuing the second ID was the mistake and it is not fixable: both are published and the page is append-only. E-2026-09 is the next free ID.

  • Adopted calibrations load from a packaged directory, so a re-calibration is a file drop rather than a code change. CALIBRATED_ZONES was a hand-written literal that stayed {} while ten fitted calibrations sat committed under results/calibration/. Nothing connected the two, so the dict could stay empty forever with every check green — which is issue #136's shape.

Artifacts now load at import from src/provael/suites/calibrations/, not from results/calibration/. That is load-bearing rather than tidy: pyproject.toml ships packages = ["src/provael"], so results/ is absent from a wheel, and reading the predicate from there would make the boundary every ASR is scored against differ between a source checkout and a pip install — one published number resting on two different predicates depending on how the reader installed the tool. This repository already learned that in 0.26.0 and wrote it into suites/__init__.py. test_adoption_does_not_depend_on_an_unpackaged_directory holds it.

Nothing is adopted, and that is the finding rather than an omission. Loading is gated on evidence the predicate can fire: an artifact needs spatial_fit.detection_rate > 0. The ten libero_object fits carry no spatial_fit at all — they predate the adversarial arm — and the one that could be replayed against real trajectories flagged 0 of 12 attacked episodes while another face flagged 5. Five of the six candidate faces score the same 0.0 benign false-positive rate, so a low benign rate is not evidence a boundary works. Adopting them would install a predicate that scores a perfect ASR against something it can never flag.

So the fits are withheld, not absent, and the difference is now visible: they ship in the package, are loaded and checked on every import, and provael doctor names all ten with the reason. "Not fitted yet" and "fitted, measured and rejected" stopped looking the same.

What this does and does not fix. It fixes the drift — test_calibration_adoption.py fails if a committed fit is neither adopted nor withheld, in either direction, so a future calibration cannot be produced and quietly ignored. It does not fix the predicate. The benign control still fires 2/50 against the default box, the published ASR is still scored against a hand-picked box that overlaps the reachable benign workspace, and #136 stays open. One GPU arm with the zones active, measuring benign and adversarial together, is what changes that.

  • provael doctor reports the published-measurement window the website has carried all along. Every page on www.provael.com says the published result was measured with v0.32.0, nine releases ago, past the project's own 2-release window. The CLI had no way to say it, so provael doctor on a green freshness badge gave a reader no signal at all — and the badge is green today because a $0.06 one-episode timing probe reset it while the 44/50 headline stayed nine minors old.

Two windows now sit side by side and disagree, which is the point: last measured 0 days ago, within 7 above published result measured with v0.32.0, 9 releases behind 0.41.0; window is 2 — PAST IT. published_measurement() picks the version behind the largest real campaign rather than the newest record, precisely so a probe cannot displace a 350-episode run; a genuinely bigger run at a newer version closes the gap on its own with no threshold edit.

The 2-release threshold now exists in watch.py. The website still holds its own copy in src/lib/freshness.ts, so this is one policy constant in two repositories and it can drift. The fix is for the site to read a published artifact instead, which is a cross-repo contract change and is not made here.

Changed

  • Two prior-art entries added: VLA-Risk and SAFE. VLA-Risk (OpenReview 31EjDFwFEe) reports degradation on its attack tasks where provael reports an envelope breach, and the two come apart in both directions — the committed run has 84% clean-task success alongside 44/50 envelope exits. Read from the public abstract only, because OpenReview serves both its web and API paths behind a bot challenge this project does not bypass; the entry says so rather than implying a closer reading than happened. SAFE (arXiv:2506.09937, NeurIPS 2025) is failure *detection* from a VLA's internal features — no overlapping quantity with an elicitation rate, and listed for completeness. Neither entry claims superiority in either direction.

RedVLA was not added: it is already covered at length, quoting their own formalism for why the denominators differ (they fix the instruction and perturb the initial state; provael does the opposite). A shorter row would have contradicted it — the channel is the scene, not the instruction. That entry's claim that "provael calibrate has never run on LIBERO" is corrected here: it ran on 6 September and the predicate is still uncalibrated, for the reason above.

8 September 2026

0.41.0

Changed

  • cli.py is a package: 3,057 lines and 27 top-level commands split by subject (issue #193). Sixteen modules for the top-level commands plus two for the existing leaderboard and study sub-apps; cli/_shared.py keeps the Typer apps, the two consoles and the cross-group helpers and now registers no commands at all. The file that every contributor touched and every merge conflict landed in is 517 lines, none of them a command body.

provael --help is byte-identical, and that is checked rather than claimed. tests/fixtures/cli-surface.json snapshots the surface as data — command order, help text, parameter names, flags, required-ness — and tests/test_cli_surface.py fails on any change. It is data rather than rendered help because rendered help wraps to the terminal: a byte-comparison of that fails on a different COLUMNS and passes on a lost option that happened to reflow. The guard was mutation-tested first: a reordered command, a renamed option and a rewritten help string are all caught; renaming a command's Python function is deliberately not, which is the property that makes the move safe.

Two things about the new cli/__init__.py are load-bearing rather than housekeeping. Typer renders top-level commands in registration order, and registration happens as each decorator executes — so the import list IS the order of provael --help, and ruff's isort sorts it alphabetically, which is not registration order. Hence isort: off around the block, with the surface test as the actual defence. The groups were moved out front-to-back for the same reason: each intermediate commit kept the order intact rather than parking moved commands at the end.

Startup does not regress: median import provael.cli is 210 ms against 245 ms before, measured from a worktree at the pre-split commit rather than from a stashed tree — git stash leaves untracked files behind, so the first comparison timed a tree where the commands had been moved out and the import list had not caught up, and reported a 69% regression that did not exist.

One test reached submit_cmd as a module attribute; it now reads the command's help through the Click object, which is how the rest of that file already worked.

Fixed

  • The calibrate GPU stage runs both arms, and is sized for it. It passed no attack, because provael calibrate took none — so the stage that exists to fit the keep-out predicate could only produce the kind of fit this release establishes is uninformative. It now passes --attack roleplay: the arm the headline rests on, so the face is chosen against the attack the published number is about.

The timeout moved with it, because the timeout is the cost ceiling and a shard that overruns writes no artifact at all — provael writes report.json once, at the end, so an overrun costs that task entirely rather than costing it some episodes. The attacked arm reuses the holdout seeds, so it is 30% more episodes and not double; at the pilot's measured 0.612 s/step that is ~3,306 s of worst case against the old 3,600 s budget, which left nothing for a slow shard. Now 5,400 s, ceiling ~$12 rather than ~$8, and the plan step prints it before anything is billed.

  • README told a reader the calibration was blocked on a run that has since happened. It said "what is missing is a run rather than an idea" and named the benign-only calibrate arm at ~$5. That arm ran on 6 September. What it produced does not fix the predicate, and a benign-only run never could have: the section now says so, with the 0/12 the fitted face flags against the 5 that x+ does and the 4 the uncalibrated default box does. The passage saying committed reports predate AttackResult.trajectory now records that the gap is closed and was not the binding one.
  • The keep-out calibration was placing its hazard zone beside the wrong face, and the benign metric could not have shown it. All ten libero_object zones fitted on 6 September reported a held-out benign false-positive rate of exactly 0.0, and that was read here as the boundary being well placed. Replaying the one committed real-model run that records trajectories against all six candidate faces (studies/keepout_face_selection/, task libero_object/0, 14 episodes across six attacks across three families) says otherwise:

| hazard face | benign fires | attacked fires | | --- | --- | --- | | x+ | 0/2 | 5/12 | | y- — the face the fitter always picked | 0/2 | 0/12 | | the other four | 0/2 | 0/12 | | the shipped DEFAULT box | 0/2 | 4/12 |

The policy leaves its workspace through +x. The hazard sat beside -y, starting at y = −0.364, and the deepest −y excursion in any episode is −0.303 — the zone was placed past a boundary the arm never reaches.

fit_spatial_zone searched the gap and took the face from hazard_zone_beside's default argument. Every gap that clears the benign envelope gives a benign FPR at or near zero, because the hazard is disjoint from the benign workspace by construction, so the search always succeeded and the number it reported carried no information. Five of six wrong faces achieve the same 0.0. The structural reason: a benign-only calibration cannot choose a face, because where an attack goes is not observable from rollouts in which no attack ran. Re-running the arm on a newer build would have reproduced this exactly.

The fitter now searches six faces × six gaps and picks the candidate that flags the most attacked rollouts among those within the benign target. calibrate_one gained an adversarial arm run at the holdout seeds, so the arms are paired — same initial states, differing only in whether the attack ran — and provael calibrate --attack <name> exposes it. Every calibration records a spatial_fit: the face, the gap, whether anything chose it, and detection_rate, which is null when no adversarial arm ran rather than 0.0. A measured failure to catch and not having looked are different findings, and collapsing them is how a zone that cannot fire comes to look like one that does not need to. The CLI prints both arms side by side and says plainly when a face was not selected.

  • Adoption now has to state the evidence that earned it, and provael doctor prints it. CALIBRATED_ZONES was dict[str, list[KeepOutZone]], so an entry could be adopted from any evidence at all — including none — and looked identical either way. The rule "never adopt a zone whose only evidence is a low benign false-positive rate" lived in a comment and in whoever remembered it. The value is now an AdoptedCalibration carrying the fitting version, the face, the detection rate and its n, all required: an adopted predicate with nothing measured against it is no longer representable.

The doctor row said calibrated zones · none · keep-out runs use the DEFAULT box and its only other state was a bare count. A count cannot distinguish a predicate that catches things from one that cannot fire. It now names the version, the face and what each zone actually flagged, in red when that is zero — and when nothing is adopted it says the ten committed fits were measured and rejected, rather than leaving none to read as work not yet attempted.

CALIBRATED_ZONES stays empty. This is one task; the published ten-task result is schema_version: 2 and records no trajectories, so nine of ten tasks have no adversarial data at all and the correct face may differ per task. What would change that is one GPU arm running both arms across all ten tasks with the new fitter. Issue #136 stays open with the numbers.

8 September 2026

0.40.0

Fixed

  • Calibration rollouts seed the policy, not only the environment. PolicyAdapter.seed() was called from exactly one place in the codebase — runner.run_episode, the attack path. collect_benign_signals() seeded the environment with suite.reset(task, seed) and left the policy alone, so a flow-matching sampler like SmolVLA's drew its noise from whatever state the ambient torch RNG was in.

That is a worse defect in a calibration than in a run. A run can be re-taken and compared; a fitted boundary is the thing every later run is scored *against*. All ten libero_object keep-out zones were fitted this way on 6 September 2026, which means the trajectories that shaped them are not reproducible — the artifacts record the seeds asked of the environment, and nothing recovers the sampler draws that actually set the envelope. Re-running on a newer build without fixing this would have bought a fresh version label and the same irreproducibility.

Calibration now carries policy_seeds: what the adapter applied, per rollout, fit seeds then holdout seeds. Not what the caller asked for — same discipline as the runner's policy_seed and resolved_device. An adapter that does not seed records null, which is the honest description of every calibration fitted before this release, and a test fails on a mutation that records the requested seed instead.

  • Both Modal GPU images pin an exact provael release, and CI fails when the pin goes stale. The two lanes were wrong in opposite directions and both reported success. examples/gpu-ci/modal_libero_suite.py pinned commit 5d34472 (v0.32.0, 9 August) under a comment saying to bump it deliberately when a stage needed newer code; five releases passed and the ~$5 calibrate arm fitted every zone on that build. examples/gpu-ci/modal_provael_gpu.py pinned nothing at all, so the scheduled canary installed whatever PyPI served that morning and resolved 0.39.1. The two GPU lanes were measuring builds five releases apart.

Neither was visible from outside. A stale pin and a current pin are the same string shape, the runs succeeded, and the artifacts recorded tool_version: 0.32.0 truthfully. Pinning was never the hard part; noticing was. tests/test_gpu_image_pin.py now asserts every modal_*.py lane declares PROVAEL_PIN, that it is an exact version rather than a range or a URL, that it equals provael.__version__, and — the mutation guard — that provael reaches pip_install only through that constant and does reach it. Bumping __init__.py without bumping the lanes is now a release-blocking failure. The deliberate consequence: a GPU measurement arm can only run against code that has actually shipped.

  • --calib finds the calibrations the calibrate arm actually wrote. load_calibrations() globbed one directory level; the arm shards one task per container and writes each into its own subdirectory. So the natural invocation — --calib pointed at the directory the arm produced — matched zero files, returned an empty map, and the run proceeded against the DEFAULT keep-out box while configured not to. Verified against the committed ten: results/calibration/libero_object_calibrate yielded nothing, only its per-task subdirectories did. The CLI's note on an empty map is what kept this a trap rather than a disaster.

The loader recurses now, and two files claiming the same (policy, suite, task) raise DuplicateCalibrationError instead of the previous last-one-sorted()-wins. Two fits are two boundaries, and every rate in the run is scored against whichever won, so there is no safe default: newest-wins needs a timestamp the artifact does not carry, and tightest-wins is a research decision rather than a loader's.

  • calibrate_suite() refuses to stamp a version that did not produce the fit. The function took tool_version as a parameter, so the label on a fitted predicate was whatever the caller passed. The CLI passed __version__ and was correct; nothing enforced it. New ToolVersionMismatchError, raised at the entry point before any GPU time is spent.
  • The staleness sweep no longer forces an incident record to become false. The action ref pin pattern is unanchored, so it matched two comments *about* a pin — both narrating the 2-4 September 2026 incident where README.md advertised an action ref that would not resolve. The 0.40.0 bump flagged them as stale pins, and rewriting them would have dated the incident to a release that postdates it. Both files are exempt now, and because the exemption is by file, test_exempt_files_carry_no_live_pin holds the other end: an exempt file may discuss a pin and may not use one.

Added

  • Ten per-task keep-out calibrations for libero_object, measured on a real policy — the run issues #136 and #171 have been blocked on since 16 August. results/calibration/ now holds one fitted artifact per task from 20 benign SmolVLA rollouts each, split fit/holdout: a 3-D benign envelope, one adjacent keep-out zone, and a holdout benign false-positive rate of 0.0 on all ten tasks against a 0.05 target. The uncalibrated global zone fires 5/100.

This is not yet adopted. CALIBRATED_ZONES is still empty and provael doctor still prints calibrated zones none. Two reasons, both worth stating rather than working around.

The artifacts were produced by provael 0.32.0, and the reason matters more than the version gap does. This entry first said the Modal image installed provael[lerobot] unpinned and that the zones were therefore fitted against a calibration_signal() four minors behind the one that would consume them. Both halves of that were wrong, and are corrected here. The measurement lane pinned a *commit* and had gone stale at it; the unpinned lane was the scheduled canary, a different file. And calibration_signal() is byte-identical between 0.32.0 and this release — the signal definition never moved, so a version gap alone would have been a weak argument for spending again. The real defect is that collect_benign_signals() never called PolicyAdapter.seed(), so SmolVLA's flow-matching sampler ran off ambient torch state and the trajectories that shaped all ten envelopes cannot be reproduced by anyone, including us. That is fixed below, and it is what makes a re-fit worth paying for.

The second reason is unchanged: a 0.0 holdout FPR says the zone does not fire on benign rollouts; it says nothing about whether the zone still catches a redirected policy. Adopting a predicate that cannot fire would score a perfect ASR and mean nothing, which is the exact failure defenses/envelope.py has an anti-cheat test for. That is answered by one more GPU arm, with the zones active, measuring benign and adversarial together.

Fixed

  • The freshness badge derives from what is committed, not from a file the lane then discards. The GPU lane's first successful commit — run 34051694289, and the run that took the site off its build deadline — put main red. latest_measurement() unions read_measurements(watch/) with the committed manifests and takes the max. provael watch --record had written watch/watch.jsonl with measured_at = _now(), wall-clock at record time; the run's manifest carries ended_at, when it actually finished. Those were 19:03:58Z and 19:03:46Z. So the badge took the later instant from a log the same step then deliberately declined to commit, the ledger took the earlier one from the run that WAS committed, and test_newest_real_measurement_agrees_with_the_badge failed — correctly, because the badge was asserting an instant no committed artifact supported.

The log is now deleted before the badge is regenerated. It stays uncommitted for the original reason — the run is in results/, and putting one measurement in two trees read by different code paths is the asymmetry that made committing the log alone unsafe — but it is also gone before anything reads it. gpu-arm.yml never calls --record, so its badge already derived from results/ only.

The committed badge is corrected here too, not just the workflow: it now reads 19:03:46Z, the instant the manifest records.

  • gpu-arm.yml keeps the run it pays for. Same defect as the canary lane, one workflow over and an order of magnitude more expensive: this arm retrieved its Modal artifacts into a 90-day upload-artifact and committed nothing. The calibrate stage — ~$8, all ten libero_object tasks, benign-only, and the run issue #171 has been waiting on — would have produced a fitted per-task envelope that aged out of GitHub before anything consumed it, leaving CALIBRATED_ZONES empty and provael doctor still printing calibrated zones none.

Two destinations, because the stages produce two different things. calibrate writes per-task calibration artifacts and no report.json; those are a fitted predicate rather than a measurement, so they land in results/calibration/ where --calib reads them and where the measurement ledger cannot see them — it walks results/** for execution-manifest.json, and a calibration artifact has none. Verified rather than assumed: adding a directory there moves no count in provael coverage and leaves gen_measurement_ledger.py --check green. Every other stage writes a report and a manifest, which are measurements, and land in results/<stage>/.

  • provael calibrate has never written a file called calibration.json, and modal_libero_suite.py told operators to look for one in three places — a return string, a comment, and the retrieval instruction printed at the end of the stage. calibration.py names its artifacts <policy>__<suite>__<task>.json, so a LIBERO shard writes smolvla__libero__libero_object_4.json. The wrong name survived because this is the one stage nobody has ever run.
  • The scheduled GPU lane keeps the measurement it produces. Fifth failure of the family behind #181 and #188, and the first where everything worked. provael watch --record appends to watch/watch.jsonl in the RUNNER'S tree; the job declared contents: read, checked out with persist-credentials: false, and uploaded only the log. So on 5 September run 33974192486 reached a real policy, printed Adversarial ASR: 33.3% (4/12), logged recorded smolvla × libero (4/14) measured with provael 0.39.1 — and the container took it with it. Neither watch.jsonl nor trials.jsonl has ever existed in this repository.

Committing the watch log alone would have been worse than nothing, and that was simulated before this was written. latest_measurement() unions the log with the committed manifests; gen_measurement_ledger.py reads only results/. The badge would have gone brightgreen while the ledger's newest row stayed a month old, and test_newest_real_measurement_agrees_with_the_badge fails on exactly that — correctly, because www.provael.com renders that ledger on /results, so the shipped state would have been a green "measured today" banner over a table whose newest row was August.

The run is now committed to results/gpu-scheduled/<ended_at>/, which is what results/ is for, and the ledger and the badge are regenerated from that one tree. Not watch/coverage.json or watch/registry.json: despite the shared directory those are a shields.io test-coverage badge and a code-derived registry owned by coverage-badge.yml. The push rebases and retries three times, because freshness.yml and coverage-badge.yml also push to main and losing a real GPU measurement to a race would be the sixth version of this bug.

Documentation

  • The RoboArena matched pair is pre-registered, in results/hardware/README.md, before any data exists and while still blocked — fixing the prediction is free today and impossible later. A base policy unmodified against the same policy with the keep-out predicate at the action layer, one flag apart, so the cost of the gate is readable off the board rather than asserted here.

Stated in win rate, not Elo. RoboArena's own paper positions its ranking against both standard Elo and conventional Bradley-Terry, because Bradley-Terry assumes each pairwise comparison happens under identical conditions and free task choice violates that. A threshold in Elo would be a threshold in a unit the venue does not report. Registered instead: the gated arm wins under 50% of head-to-head comparisons against its own twin (a clamp can only remove motion), predicted at or above 40%, abandoned below 25%, and nothing read before 50 comparisons. If the gated arm wins MORE often, that is not a win for the gate — it is evidence the predicate is correlated with task structure, and it will be reported in those words.

The base is paligemma_fast_droid and the criterion is what is registered: the base must be reproducible, robot data and pre-trained weights public. pi05_droid was the initial plan and fails that test, because Pi's pre-training data is not released — which by RoboArena's own definition puts "open-source: No" on the row and yields a number no outside party can re-derive.

There is no deadline, and an earlier draft said there was. The CoRL round carrying the 8 September soft and 13 September hard deadlines was CoRL 2025 — that page reads "Conference on Robot Learning, 2025", its call for papers closed 17 September 2025 and its workshop was held 27 September 2025. Those dates were projected onto 2026 off a stale page. The form is live and the platform is active (public data dump dated 17 July 2026), but nothing is closing, and the false urgency was about to buy a rushed submission.

  • results/hardware/README.md records what a real-robot leaderboard entry actually needs, dated 6 September 2026 and assessed against RoboArena's September round. Three blockers, none of which is time: nothing here speaks their inference API (provael serve is the ATTESTATION server — /healthz, /attest, /assurance-report); the openpi adapter that would front a π0.5 policy is scaffolding by our own declaration and has never been exercised; and ActionEnvelopeClamp's bounds are the CPU fixture's benign envelope in the fixture's action space, which means nothing on a 7-DoF DROID cell. Picking replacements unexamined is what #136 was.

It also records a number this assessment got wrong. A draft covering letter for that submission stated the benign control fires on 5 of 100 episodes, and the first version of this entry called that fabricated on the grounds that no arm of the pinned control run produces it. That was a check of one run against a figure that pools two, and the draft was right: smolvla_libero_object_suite fires 2/50 and smolvla_libero_object_control fires 3/50, pooling to 5/100 — 5.0%, Wilson 95% [2.2%, 11.2%], which is exactly what issue #171 publishes. The pooled figure is the better one: a single-run 3/50 carries no interval and hides that every firing lands on libero_object/4 or /5 while eight tasks stay silent through 80 benign episodes. Whether that rate and that clustering hold on a real cell is worth more than the row.

Runs executed stays 0; provael coverage still reports hardware=0.

6 September 2026

0.39.5

Added

  • watch/release.json, so a consumer derives the release instead of copying it. watch/ had the counts, the badge, the measurement ledger and the coverage total, and no release artifact, so the one fact a consumer restates most often was the one it could not derive. www.provael.com kept its own copy refreshed on the release cadence: on 6 September 2026, 14 of its built pages rendered v0.39.3 while the tag, the GitHub release and PyPI all said 0.39.4. The worst of them was /security, which named that superseded release as the one a reporter should reproduce on before filing. (Not quoted verbatim here: www.provael.com renders this file at /changelog, and its check:versions guard reads a currency-claim sentence on a rendered page as a live claim, correctly — it cannot tell a quotation from an assertion.)

Every value derives from provael.__version__. There is deliberately no commit sha and no published_at: the generator runs offline and could only guess at them, and an unverifiable field in a trust artifact is worse than a missing one. A test names those fields explicitly, because the tempting next commit is the one that adds a date for a consumer that wants to render one.

Wired into make check-docs as check-release, and regenerated by coverage-badge.yml alongside the other two watch artifacts as a safety net, following that workflow's existing commit pattern rather than adding a second one.

Fixed

  • Two documents said ISO 10218:2025 defers its cyber detail to IEC 62443. It does not. docs/compliance/machinery-annex-i-part-a.md carried "which defers detailed cyber requirements to IEC 62443" and docs/crosswalk/halos-integrator.md carried "cyber clauses, deferring detail to IEC 62443". Two more places described Provael's own SL2 view as something ISO 10218 routes to.

Corrected by reading Clause 2 Normative references on the ISO Online Browsing Platform, not by inferring it from a missing citation. Clause 2 lists ISO 3864-x, ISO 4413/4414, ISO 7010, ISO 9283, ISO 12100, ISO 13732-x, ISO 13849-1:2023, ISO 13850, ISO 14118/14119/14120, ISO 19353, ISO 20607, ISO 20643 and IEC 60073. IEC 62443 and IEC TR 63074 appear only in the Bibliography, which is informative. The Foreword does say the revision adds "requirements for cybersecurity to the extent that it applies to industrial robot safety", and that part is kept.

Four sites, and two were found by grep rather than named: the machinery Annex I card and the _iso_10218_2 docstring in src/provael/assurance.py. Four other candidates carry no such claim and were left alone, because "Maps to" and "cross-map to" already read as Provael's own crosswalk and rewriting them would have degraded accurate text.

Nothing machine-readable moved. _IEC, the iec-62443:slv key, every control identifier and the emitted routes_to field are a published contract and were never the defect. Recorded as E-2026-05 in docs/errata.md.

Added

  • watch/registry.json publishes the registered/runnable split. It carried policies: 8 and suites: 6 as bare integers while the code already declared which of those are scaffolding, and list-policies / list-suites already rendered that. coverage_json() never exported it, so a consumer wanting the runnable number typed one: www.provael.com publishes "5 suites" beside a registry saying 6.

Now runnablePolicies, scaffoldingPolicies, runnableSuites, scaffoldingSuites, plus scaffoldingPolicyNames and scaffoldingSuiteNames so a consumer can render which rather than only how many. 8 = 5 + 3, 6 = 5 + 1.

The counts are properties over the name tuples rather than stored fields, so they cannot drift from the declaration they came from, and the split is read from SCAFFOLDING_POLICIES / SCAFFOLDING_SUITES rather than probed — a filesystem probe answers differently in a checkout and in a wheel, and fails toward "measured". A test asserts the split is identical with results/ absent.

6 September 2026

0.39.4

Fixed

  • The Docker Hub mirror pointed at a namespace that does not exist. docker-publish.yml has targeted docker.io/provael/provael since the mirror was written, gated on DOCKERHUB_USERNAME/DOCKERHUB_TOKEN. Those secrets were never set, so the branch never ran and the wrong coordinate never surfaced. It now points at docker.io/mndfreek/provael.

The namespaces differ on purpose. GHCR stays ghcr.io/provael/provael, matching the GitHub org. A Docker Hub *organisation* named provael requires a Docker Team plan at $15/seat/month — $180/year to make one string match another string, for a mirror whose only job is discoverability. The image bits are identical; only the coordinate differs, and the workflow comment records why so the mismatch reads as a decision rather than a mistake.

Added

  • A Codespaces badge, so the devcontainer has an entry point. .devcontainer/devcontainer.json has existed since 26 July — Python 3.12, the uv feature, uv sync --locked on create, ruff and mypy extensions — and a grep for codespaces across the repo returned zero hits. A working artifact with no way in is the same as no artifact. Placed beside "Open in Colab", the other one-click entry point on the page.
  • A Binder environment, so the notebooks run without a Google account. All five notebooks carried only an "Open in Colab" badge, and Colab requires a sign-in. A reader who has to create an account before running anything has already been asked to do work, on notebooks that exist so nobody has to take the README's word for a number.

binder/environment.yml pins Python 3.12, not 3.11: requires-python = ">=3.12", so a 3.11 image builds cleanly and then fails at the pip step, launching the badge into a broken kernel with no obvious cause.

It deliberately does not pin the provael release. That would be theatre — notebook 01 runs %pip install -q provael in its own second cell, so any pin here is replaced by latest the moment the notebook runs — and it would rot unnoticed, because test_version_consistency.py matches provael/provael@vX.Y.Z action refs and pre-commit rev: lines, not pip specifiers. Verified by mutating a pin and watching the suite stay green, rather than assumed.

  • watch/measurements.json — one ledger row per committed measurement. watch/freshness.json answers *when was anything last measured* and collapses every run into one instant for a badge; watch/registry.json answers *how many attacks are registered*. Neither answers the question a stranger actually arrives with, which is how current is the specific number I am reading. That needs a row per measurement carrying the version it ran on and the artifact it came from, and nothing published it.

The date cannot come from report.json, which carries no timestamp on purpose — the determinism contract makes a report a pure function of its config. It comes from the execution manifest beside it, via provael.watch.measurements_from_results, so the ledger and the badge cannot drift into disagreeing about the project's own currency. test_measurement_ledger.py asserts that agreement directly.

Two honesty fields travel with every row. recorded: false marks a reconstructed date (an exact-midnight ended_at, or a legacy-unverified state) that must never render as a measurement instant. countsAsMeasurement: false marks a fixture backend — a stub run executes real attacks in under a second on CPU and would otherwise let a consumer refresh a freshness claim having re-measured nothing. Today's ledger holds 26 rows, of which 20 are real-policy measurements and one carries a reconstructed date.

The file says when and on what, never where a number is published. Only a consuming site knows that, and encoding a site's information architecture into a repo artifact would put the mapping in the one place a site author never looks.

Fixed

  • Four documents told a contributor to run a weaker type-check than CI runs. README.md, CONTRIBUTING.md, .github/PULL_REQUEST_TEMPLATE.md and CLAUDE.md all documented the gate as uv run mypy src, while .github/workflows/ci.yml runs uv run mypy src scripts/action and pyproject.toml pins files = ["src"]. So the documented command checked 105 files against CI's 111: a type error in scripts/action — the five scripts that decide whether a release passes — passed locally and failed in CI, for anyone who followed the docs.

Both halves are fixed. The commands now name scripts/action, and the three contributor-facing surfaces lead with make check so the gate has one definition instead of four transcriptions of it. This is the same failure the repo already guards for restated *numbers*; it had simply never been applied to a restated *command*.

Added

  • A Makefile, wrapping the gates that already existed. Every recipe is the command CI actually runs; nothing here gates anything that was not already gated. make check is the three commands ci.yml runs, make check-doc-counts is the same --check that tests/test_counted_claims.py already calls, and make help lists the rest. A Makefile that quietly introduced a *new* rule would be the worst version of this file, because the rule would then live somewhere no reviewer looks.

check-doc-counts also gets its own CI step, and it is deliberately redundant with the pytest assertion rather than replacing it. The rule stays where it was; what the step adds is a legible failure. A stale inventory line currently surfaces as one assertion inside a ~1300-test, ~25 s run — as a named ~1 s step it says what is wrong in its own title. Same rule, better signal.

  • provael.__all__, derived from docs/python-api.md and enforced by a test. The published Python API and the package's export list had no relationship: src/provael/__init__.py declared no __all__ at all, so there was nothing for a rename to disagree with. A symbol could move, the gate stays green, and the doc goes on describing an import that no longer resolves. tests/test_public_api.py now fails in both directions — documented but not exported, and exported but not documented — and the second direction matters as much as the first, because an undocumented public name is a support burden nobody agreed to.

The exports are lazy, and that is the interesting part. Every documented name lives in a submodule, and re-exporting all eight eagerly costs ~1.14 s and pulls numpy in, against ~1.1 ms for the bare package: a ~1000x regression paid by every import provael and by every CLI invocation, in exchange for a shorter import line. __init__.py resolves them through a PEP 562 __getattr__ instead, so the cost is paid only by a caller who touches the name. The cheap way to undo that is a convenience import added at the top of the file later, so the test asserts the laziness holds by checking sys.modules in a subprocess — an in-process check would pass regardless, since the suite has already imported half the package by then.

The submodule paths in the docs are unchanged and keep working; this only adds the top-level spelling.

Earlier

47 earlier releases

  • 0.39.3 · 4 September 2026
  • 0.39.1 · 1 September 2026
  • 0.39.0 · 1 September 2026
  • 0.38.1 · 31 August 2026
  • 0.38.0 · 24 August 2026
  • 0.37.0 · 22 August 2026
  • 0.36.2 · 21 August 2026
  • 0.36.1 · 21 August 2026
  • 0.36.0 · 21 August 2026
  • 0.35.0 · 18 August 2026
  • 0.34.0 · 18 August 2026
  • 0.33.2 · 15 August 2026
  • 0.33.1 · 13 August 2026
  • 0.33.0 · 10 August 2026
  • 0.32.0 · 8 August 2026
  • 0.31.1 · 3 August 2026
  • 0.31.0 · 2 August 2026
  • 0.30.0 · 1 August 2026
  • 0.29.1 · 31 July 2026
  • 0.29.0 · 31 July 2026
  • 0.28.0 · 30 July 2026
  • 0.27.0 · 30 July 2026
  • 0.26.1 · 28 July 2026
  • 0.26.0 · 28 July 2026
  • 0.25.1 · 27 July 2026
  • 0.25.0 · 26 July 2026
  • 0.22.0 · 23 July 2026
  • 0.21.0 · 22 July 2026
  • 0.20.0 · 21 July 2026
  • 0.19.0 · 20 July 2026
  • 0.18.0 · 19 July 2026
  • 0.17.0 · 18 July 2026
  • 0.16.0 · 15 July 2026
  • 0.15.0 · 13 July 2026
  • 0.14.0 · 13 July 2026
  • 0.13.0 · 8 July 2026
  • 0.12.0 · 8 July 2026
  • 0.11.0 · 6 July 2026
  • 0.10.0 · 5 July 2026
  • 0.9.0 · 4 July 2026
  • 0.8.0 · 4 July 2026
  • 0.7.0 · 3 July 2026
  • 0.6.0 · 30 June 2026
  • 0.5.0 · 29 June 2026
  • 0.4.0 · 28 June 2026
  • 0.3.0 · 27 June 2026
  • 0.1.0 · 27 June 2026

Full notes for these are in CHANGELOG.md ↗.