
OpenAI says coding agents now outpace human labor time inside its research group
OpenAI says coding agents now outpace human labor time inside its research group, while safety controls still shape model work.
OpenAI published a new look inside its research organization on September 6, saying coding agents have moved from occasional assistance to a daily part of frontier model work. The company says the median researcher was using more than $600 per day of agent inference at API prices by mid August, while the 90th percentile user exceeded $7,000 per day.
The figures matter because they describe how AI labs are trying to speed up the work that produces stronger models. OpenAI says total agent runtime across its research organization has now passed total human labor when converted into standard eight hour workdays, reaching 3.1 agent workdays for every human workday. That does not mean research output rises at the same rate. The company directly cautions that AI research has bottlenecks in judgment, compute, safety review and deployment decisions.
What Changed
The update says researchers are using agents in more concurrent workflows and assigning them a broader range of tasks. OpenAI says coding agents are still strongest around code, infrastructure help, monitoring runs and experiment support, while high level planning remains a small share of agent output tokens. The company also says researchers are writing more code and running more experiments in 2026, with August reaching a high point since its tracking began in January 2025.
OpenAI also disclosed a practical constraint. After a recent incident involving agents and research infrastructure, it temporarily shut down a container service used for training, then restored it with added restrictions. Later, preliminary evidence that Astra may have critical cyber capabilities led to tighter model specific security controls. OpenAI says Astra class GPU allocation fell 59.2 percent in the following week, while allocation to other model classes rose 17.2 percent.
Why It Matters
For developers and enterprise buyers, the useful signal is not that coding agents replace research teams. It is that advanced users are building workflows around many assisted sessions at once, and they still need human review when task complexity rises. OpenAI says over half of successful four to eight hour tasks in the last six months required at least one intervention.
- Agent adoption can increase experiment throughput, but reliability still depends on oversight.
- Security controls can redirect compute rather than simply slowing all work.
- Teams evaluating coding agents should measure intervention rate, not only task completion.
The CyberOGZ read is that internal agent productivity metrics are becoming strategic disclosure, especially when they touch model pace and safety. Readers should watch whether other labs publish comparable measurements, because shared definitions would make claims about AI accelerated research easier to compare and harder to inflate.
Sources
Cover photo by Ludovic Delot on Pexels, used under the Pexels License.
CyberOGZ Team






Comments (0)
Leave a Comment