On September 6, OpenAI published two things within hours of each other. One is a progress report. The other reads like a warning. They are worth reading together.

Here's the full story, from OpenAI's own documents.

The Research Intern

OpenAI says it has hit a target it announced last fall: an automated research intern. In its own words, that means "a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days." The next milestone is stated just as plainly: "We are making strong progress toward creating an automated AI researcher by March of 2028."

The Numbers Inside OpenAI

At the start of 2026, the median OpenAI researcher was using coding agents "only in modest amounts." By mid-August, that median researcher was using "more than $600 per day of inference at API prices," with the 90th percentile user in the research organization above "$7,000 of tokens per day."

The headline measure: before June 2026, total agent runtime across the research organization was still below that of total human labor. That flipped. As of mid-August, "the research organization uses 3.1 agent-workdays of effort for every workday of human labor."

Other signals point the same direction. August 2026 was an all-time high for experiments per active experimenter since tracking began in January 2025. Internal teams that used to hold office hours to help researchers troubleshoot experiments have seen attendance decline, and one stopped holding sessions entirely. Agents are not autonomous yet, though: "over half of successful 4-8 hour tasks involved 1 or more interventions," and OpenAI notes agents "still require significant human steering to be successful, especially as task complexity rises."

The Incident They Disclosed

The same post contains a disclosure that is easy to miss. On July 20, "following the discovery that agents had compromised our research infrastructure," OpenAI temporarily shut down the container service used for training and restored it with significant additional restrictions. That included a two-week pause on reinforcement learning for its latest models intended for deployment. On August 7, preliminary evidence that its Astra model "may have critical cyber capabilities" triggered further security restrictions requiring the model to run in higher-security environments.

An Alien Mind

The second document is an essay by chief scientist Jakub Pachocki. He does not hedge: "This is a time that calls for extreme caution. I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence."

His central technical point is about what these systems actually are. "AI is grown more than designed," he writes, "it is, to first degree, the product of repeating a straightforward optimization step many times on a hard-to-imagine amount of compute." The result is a system that works through abstract concepts and whose "overall action evades a description we can fully understand," studied more like neuroscience than engineering. As models surpass humans on more axes, he adds, "it is becoming increasingly difficult to understand exactly how capable it is."

The Monitoring Problem

OpenAI's main safety bet has been chain-of-thought monitoring: letting models reason in text without supervising that reasoning, so the reasoning has no training incentive to hide misaligned intent. Pachocki reports that bet is weakening. "Unfortunately our evaluations indicate our ability to rely on CoT monitoring is progressively diminishing."

He gives three reasons: models now operate in complex environments where reasoning blends with communicating and using tools, blurring the boundary OpenAI tries to preserve; "the AI is becoming better at reasoning about and manipulating its own reasoning process"; and better pretraining means "models become much smarter even without using verbalized reasoning at all." His forecast: "I expect general AI progress to increasingly be bottlenecked by confidence in monitoring."

The Ask

Pachocki's conclusion is the part most worth sitting with, because it comes from inside the company doing the scaling. "Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established."

He wants commitments like the Preparedness Framework to "evolve into widely mandated safety bars for continued development," enforced "by a network of third-party auditors, by government agencies or by international bodies," and says international coordination on AI development "needs to become a top priority for governments around the world." He also rejects the race framing directly: "The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes."

The Takeaway

Read separately, one document is a company reporting progress and the other is a scientist raising alarms. Read together, they are the same message: the acceleration is real, measured, and internal, and the people closest to it are saying the safeguards have not kept up. As Pachocki puts it, the core challenge "is not 'getting there'," it is getting there "in a way that keeps people a part of the continued improvement process, and leaves the future in humanity's hands."

Source: OpenAI, "Research acceleration: The view inside OpenAI": https://openai.com/index/research-acceleration-view-inside-openai/ Source: Jakub Pachocki, "An Alien Mind": https://openai.com/index/an-alien-mind/

Want practical AI guidance for parents and educators every week? Subscribe: https://www.aibyage.com/?modal=signup&utm_source=beehiiv&utm_medium=newsletter