
AWS says Continuum reached 89% on end to end code vulnerability benchmark
AWS says Continuum hit 89.0% on CyberGym-E2E, showing progress in AI systems that find, prove and patch code vulnerabilities.
AWS says its Continuum system for code vulnerabilities reached an 89.0% end to end success rate on CyberGym-E2E, a public benchmark that asks an agent to find a flaw in a real codebase, demonstrate it with a proof of concept, then repair it without breaking the project tests. The company says Continuum passed 819 of 920 tasks within the benchmark's 90 minute limit, beating the previous public high of 65.9% by 23.1 percentage points.
The announcement matters because it moves the security AI conversation past simple bug spotting. Many developer tools can flag suspicious code. The harder workflow is proving that a flaw is exploitable, creating a patch, and checking that the fix does not damage expected behavior. AWS is pitching Continuum as a system for that full loop, not only as another scanner that leaves engineers to sort a long queue of alerts.
What AWS measured
CyberGym-E2E contains 920 tasks based on historical OSS-Fuzz vulnerabilities across 139 open source projects. AWS says the median project has more than 600,000 lines of code, and each task runs in a container with the tools needed to build and test the project. The agent receives no vulnerability description, crash log, proof of concept, or original patch. External network access is blocked during the run.
The benchmark scores cumulative stages. The system first has to produce an input that crashes the vulnerable program. It then has to patch the code so that crash no longer works. Next, the patched project must still pass its functionality tests. A final diagnostic stage checks whether the patch also fixes the selected historical vulnerability. AWS says Continuum reached 92.5% at the crash reproduction stage, 89.6% at repair, 89.0% after functionality tests, and 37.8% on the selected historical target.
Why teams should read the result carefully
The useful CyberOGZ takeaway is that autonomous security tools are becoming more relevant to the middle of vulnerability work, where teams spend time reproducing bugs, validating priority, and checking repairs. That could help defenders handle more code review without treating every AI finding as an emergency.
There is still a boundary around the claim. AWS says CyberGym-E2E focuses on memory safety vulnerabilities in C and C++ projects, where sanitizer crashes provide objective evidence. That does not cover many production issues involving identity, business logic, cloud configuration, web authorization, or data exposure. Buyers should treat the result as strong evidence for one demanding class of security work, then ask how the same system proves impact and avoids unsafe changes in their own repositories.
Sources
Cover photo by Pixabay on Pexels, used under the Pexels License.
CyberOGZ Team






Comments (0)
Leave a Comment