What the buyer cannot prove today.
Most agent evaluations reward successful completion. That can hide unsafe sequencing, stale evidence, excessive tool authority and agents that continue after the permission envelope has changed.
A bounded agent benchmark for tool use, changed context, revoked authority, conflicting evidence and irreversible action boundaries.
Most agent evaluations reward successful completion. That can hide unsafe sequencing, stale evidence, excessive tool authority and agents that continue after the permission envelope has changed.
• Declared agent permission envelope
• Pre-registered authority-change scenarios
• Tool-call and execution-boundary tests
• Changed-context and revocation tests
• ALLOW / HOLD / DENY / ESCALATE behavior report
• Evidence package suitable for internal assurance or publication
→ Capability is not autonomy.
→ Autonomy is not authority.
→ Task success can still be governance failure.
→ The benchmark tests whether execution remains inside the authorized boundary.
AI agent benchmark · AI agent permission testing · AI agent governance benchmark · agentic AI authorization · AI tool use safety benchmark
Live cumulative activity recorded across the public Exchange surface.