Microsoft run-assert-eval links AI agent risk discovery to runtime policy checks

Microsoft run-assert-eval links AI agent risk discovery to runtime policy checks

Microsoft's run-assert-eval links AI agent risk discovery, runtime policy and repeated evaluation in one open source workflow.

Format News Brief
Read Time 2 min
Category Cyber Security
Updated Sep 27, 2026

Microsoft has introduced run-assert-eval, an open source skill for developers who need evidence that an AI agent risk was found, controlled and tested again. The tool connects three pieces Microsoft has been building in public: Clarity for threat modeling, ASSERT for behavior based evaluations and Agent Control Specification for runtime policy.

The practical change is that a developer can describe an agent in VS Code and run one loop that discovers likely failure modes, turns selected risks into evaluation suites, generates an ACS policy and repeats the same evaluation after the control is applied. Microsoft says the point is to keep the behavior definition, test cases and judge constant so the policy is the main thing that changes between the baseline and governed runs.

Why teams should care

Agent safety work often breaks into separate handoffs. One group writes requirements, another runs tests, someone else writes a rule and the proof that the rule helped arrives later, if it arrives at all. run-assert-eval is aimed at that gap. It treats the failure measurement and the mitigation as one workflow, which matters for agents that can call tools, retrieve private records or make changes in business systems.

Microsoft's worked example uses a billing support agent tied to one customer account. In the baseline test, the agent disclosed another customer's data in 12 of 40 applicable conversations, an observed 30.0 percent rate. After applying runtime policy, the governed run showed two violations in 34 applicable conversations, or 5.9 percent. Another split showed cross-customer scenario violations falling from 43.8 percent to 0.0 percent, while permissible behavior violations fell to zero across the reported splits.

The decision value

The useful part is not the specific billing demo. It is the discipline of freezing the comparison before claiming that an agent is safer. Microsoft also notes that generated policy still needs human review, including the manifest, intervention point and target wiring. That is a healthy boundary. Teams should treat this as a release gate candidate, not a replacement for security ownership.

For developers building production agents, the takeaway is simple: a control that blocks risky tool calls is stronger when it is tied to the exact evaluation that exposed the risk. The next thing to watch is whether teams outside Microsoft can repeat these results on messier agents, larger samples and regulated workflows where false refusals carry real cost.

Sources

Cover photo by Daniil Komov on Pexels, used under the Pexels License.

Comments (0)

Leave a Comment

Loading comments...