01—OpenAI says the research intern milestone is here
OpenAI has been talking for months about a near-term milestone for autonomous research: an AI system that can take on a well-defined research task, work through it under human direction, and return with something useful after the kind of effort that might normally occupy a skilled researcher for several days.
On September 6, the company said it has reached that point.
OpenAI calls the milestone an “automated research intern.” The phrase is deliberately narrower than “automated scientist.” People still choose the research direction, decide which ideas are worth pursuing, and judge the results. The agent’s job is to take a bounded piece of research work and carry more of the execution on its own.
The more revealing part of OpenAI’s update is what has already happened inside the lab. Coding agents are no longer occasional helpers for researchers. They are running throughout the day, often several at once, and their total runtime has grown past the amount of human labor measured across the research organization.
02—3.1 agent workdays now sit behind every human workday
OpenAI’s clearest number is also the easiest one to picture.
By mid-August, the company says its research organization was using 3.1 agent-workdays of effort for every human workday, based on an eight-hour workday. Before June, total agent runtime was still below total human labor. That balance has now flipped.
This does not mean three AI researchers have literally replaced every person in the building. It is a measure of runtime: researchers are launching coding agents in parallel and letting them keep working while the human moves between experiments, reviews results, or starts another task.
That kind of concurrency changes the shape of the workday. One researcher can have several lines of investigation or engineering moving at the same time instead of waiting for one task to finish before starting the next.
OpenAI says the number of researchers using highly concurrent workflows — four or more agents running simultaneously — is increasing as well. The “fleet of agents” idea is already becoming a normal part of the internal research workflow, not a future concept.
03—The agents are doing more than writing code
Coding is still central, but OpenAI’s own breakdown shows the work widening.
The company maps agent activity across six parts of the AI research lifecycle: deciding what to work on, designing ideas and engineering specs, building code and datasets, running training and evaluations, analyzing results, and communicating findings and decisions.
All six categories grew between January and August 2026. Research and infrastructure code remains a major use case, but technical help and monitoring runs expanded noticeably as well.
OpenAI says researchers are also delegating higher-level and longer-horizon tasks more often. That is the shift behind the “research intern” label. The useful unit is no longer just “write this function” or “fix this script.” Agents are being asked to own a larger chunk of a research task and keep context across the steps needed to complete it.
High-level planning still remains a small share of agent output. Humans continue to set priorities and make the larger research decisions. The automation is growing most quickly in the execution layer underneath those decisions.
04—OpenAI researchers are running more experiments
Research speed is difficult to reduce to one metric, but OpenAI points to two places where the change is already visible: code and experiments.
The company says researchers are contributing code faster and running more experiments. Experiments per active experimenter rose through 2026, with August reaching the highest level OpenAI has measured since it began tracking the metric in January 2025.
OpenAI notes that more available compute also contributes to that increase, so it does not treat agent adoption as the only explanation. Still, the correlation with growing Codex use is strong enough that the company sees coding agents as a meaningful part of the acceleration.
This is where agent concurrency becomes especially useful. Research often involves a large amount of setup work around the central idea: writing evaluation code, building infrastructure, debugging experiments, comparing outputs, and preparing another run. Agents can take on several of those loops while the researcher keeps the larger experiment moving.
The result is not simply more generated code. It is a higher-throughput research workflow.
05—Some internal support work is already disappearing into the agent layer
One small detail in OpenAI’s report says a lot about how deeply the tools are being used.
Researchers have increasingly turned to coding agents for troubleshooting internal research infrastructure. OpenAI says multiple teams that once held office hours to help colleagues debug experiments have seen attendance decline during 2026. One team stopped holding those sessions entirely and shifted its time toward improving other systems.
That is a very different kind of automation from code completion.
The agent is not just producing an artifact. It is absorbing a support workflow that used to require another person to stop what they were doing, inspect the problem, explain the fix, and help the researcher get moving again.
For a research organization, removing that friction can compound quickly. Every blocked experiment that gets unstuck without waiting for a human support window can move back into the research loop sooner.

06—The target after the intern is a full automated AI researcher
OpenAI is treating the research-intern milestone as a step, not the end point.
The company says it is working toward an automated AI researcher by March 2028: a system that can operate under human supervision while contributing more deeply to research on model capabilities, alignment, and related work.
That timeline gives the current update more context. OpenAI is not only measuring whether coding agents can complete isolated software tasks. It is tracking how much of the research lifecycle they can absorb, how long the tasks can run, how often researchers use them concurrently, and how the surrounding organization changes as those capabilities improve.
The company also says humans remain responsible for research priorities, deciding which ideas and results to pursue, and determining when work should move forward. The near-term picture is therefore not an autonomous lab operating by itself. It is a lab where humans increasingly direct a large amount of machine-executed research work in parallel.
07—The research organization is becoming a good test of what agentic work looks like
There is a reason this internal data is more interesting than another coding benchmark.
OpenAI’s researchers are using the agents on the same infrastructure, experiment, debugging, and analysis work that contributes to new models. The tools are being tested inside a workflow where small efficiency gains can stack across thousands of tasks.
The company’s median researcher was already using coding agents daily by mid-August, with more than $600 per day of inference measured at API prices. That figure is less useful as a budget number than as a signal of intensity: agent use has become continuous enough to represent substantial compute consumption for a typical researcher.
And because researchers can run multiple sessions at once, the scaling unit is no longer one human paired with one assistant. It is one human directing several machine workers, reviewing their outputs, and deciding where to spend the next round of attention.
That may be the more durable lesson from the “research intern” milestone. The capability of an individual agent matters, but so does the workflow that lets one person direct many of them without losing control of the research.
08—The Upgrade Feeling
The headline sounds futuristic, but OpenAI’s own numbers make it feel much more immediate.
The company says it has reached the automated research-intern milestone it set for September. Inside the research organization, agent runtime has already grown to 3.1 workdays for every human workday, concurrent sessions are becoming more common, and experiment throughput is rising.
The interesting shift is organizational. Coding agents are becoming part of the machinery of research itself: writing code, debugging infrastructure, monitoring runs, and carrying longer tasks while humans stay focused on direction and judgment.
If that pattern keeps expanding, the next leap in AI research may not come from one researcher working faster. It may come from every researcher being able to direct a small, persistent team of agents at the same time.
