ProductEvidenceTop 10ComplianceDocsStar on GitHubQuickstart
TEXT LAYER VS. ACTION LAYER · PV-015

Text-layer red-teaming stops at the sentence. The robot doesn’t.

garak, PyRIT, andpromptfoo are excellent at scanning what a model says. But a vision-language-action policy turns language into motion — so the failure is a trajectory, not a paragraph. That action layer is the part text-only tools structurally can’t reach, and it’s the whole reason Provael exists. These tools aren’t rivals — they’re the other half of your stack.

What each tool actually tests · LLM red-teaming tools cover the language layer; Provael covers the action layer
ToolWhat it testsSuccess =TaxonomyEvidence
garak
NVIDIA
LLM text & prompt I/OJailbroken / toxic / leaked textOWASP LLM Top 10 · MITRE ATLASVulnerability-scan report
PyRIT
Microsoft
Generative-AI & agent I/OHarmful / policy-violating textOWASP · MITRE ATLASRed-team eval logs
promptfoo
OpenAI (acq. 2026)
LLM evals & prompt red-teamingFailed eval / jailbreakOWASP LLM Top 10Eval reports
Provael
the action layer
VLA robot policy — the commanded actionPolicy driven to an unsafe action (keep-out), as a calibrated ASREmbodied AI Security Top 10SARIF + conformity evidence (EU AI Act / Machinery Reg / ISO)

Questions people actually ask

Can garak or PyRIT test a robot policy?

They test the language and tool-call layer brilliantly — but they do not model the action space. A vision-language-action policy turns an instruction into a trajectory, and that trajectory is where the embodied failure lives. Text-only tools have nothing to score there. Use them for the language layer; use Provael for the action layer.

Is Provael a replacement for LLM red-teaming?

No — it is complementary. If your robot has an LLM in the loop, red-team the text with garak or PyRIT and red-team the resulting motion with Provael. The two cover different halves of the same system.

What makes Provael’s number different from a jailbreak demo?

Every Provael result is a calibrated attack-success rate with a 95% Wilson confidence interval and a matched benign-control false-positive rate — plus published null results when an attack does not transfer. That is a measurement you can file, not a one-off "we broke it."

Do I need a GPU?

No. Provael is CPU-first: the deterministic core and the full attack suite run without a GPU. Real-model transfer tests are optional and gated behind an extra.

Red-team the layer no one else covers.

Keep your LLM red-teaming. Add the action layer — a calibrated, reproducible attack-success rate for your robot’s policy.