OpenAI Now Runs 3.1 Agent Workdays for Every Human Workday. Over Half of Its Long Agent Tasks Still Needed a Person to Step In.
On September 6 OpenAI published its own internal agent numbers. Its research organization now uses 3.1 agent workdays for every human workday, and over half of the successful 4 to 8 hour agent tasks in the last six months required at least one human intervention. The company building these models measures its own autonomy rate and publishes it. Ask the vendor selling your store a hands-off AI BDC for the same two numbers.
Adam founded Savvy Dealer and has spent 30 years at the intersection of automotive retail and digital strategy.

Want to Learn More?
Book a quick demo to see these strategies in action.
On September 6, OpenAI did something almost no vendor selling AI into your dealership does. It published its own numbers.
The post is called Research acceleration: The view inside OpenAI, and it reports how OpenAI's own research organization actually uses coding agents day to day. It is the closest thing anyone has released to an audited account of what AI agents do when a serious company turns them loose on real work, at scale, on the workload best suited to them.
Two numbers in it matter to a car dealer.
The first is the one everybody will quote. "In terms of a standard 8 hour workday, as of mid-August, in total, the research organization uses 3.1 agent-workdays of effort for every workday of human labor." Before June 2026 that ratio was under one. In under three months, agents went from doing less total work than the humans to doing more than three times as much.
The second number will not appear in a single sales deck this year. "In the last 6 months, over half of successful 4-8 hour tasks involved 1 or more interventions."
Read that again. Those are the successes. More than half of them needed a human to reach in and steer.
What OpenAI actually claimed
The headline milestone is real and dated. "According to our measurements, we have now reached the goal, announced last fall, of having an automated research intern by September of this year."
OpenAI defines the term carefully: "a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days." The company says it is "making strong progress toward creating an automated AI researcher by March of 2028." Unite.AI covered the announcement the same day.
OpenAI also put a caveat in its own post that most coverage skipped: "AI research is a complex process with many potential bottlenecks, so the overall pace of progress likely won't keep pace with these specific metrics."
That is the company with the strongest possible incentive to oversell, telling you its own headline metrics run ahead of reality.
The word they picked was intern
Every dealer principal already knows exactly what an intern is worth.
You give an intern defined work with a defined end. Pull the aging report. Photograph the 40 units that came off lease. Call the 200 unsold service customers from last quarter and log the answers. Real work, real value, and none of it requires the intern to decide anything that costs you money if it goes wrong.
You do not hand the intern the used car desk. You do not hand them the co-op submission or the conversation with the customer who is upside down $9,000. The reason has nothing to do with the intern being slow or careless. Those jobs require judgment about consequences, and judgment is the part you keep.
OpenAI kept it too. Buried in a post whose entire purpose is to show how much agents are doing: "People still set our research priorities, judge which ideas and results to pursue, and decide whether to scale, pause, or deploy systems."
The part about money
The cost figures are worth your attention because of what they imply about how AI is priced to you.
At the start of 2026 the median OpenAI researcher used coding agents "only in modest amounts." By mid-August that median researcher was "integrating agents daily into their work, using more than $600 per day of inference at API prices." The 90th percentile user "now uses more than $7,000 of tokens per day."
Now look at the flat monthly seat price on the AI tool your store is being pitched. Either it is doing a far smaller amount of work than that, which is fine as long as you know it, or the vendor is absorbing a metered cost that will get repriced at renewal. Both are survivable. Being surprised by either one is not.
Ask what the usage cap is and what happens when you hit it. If the answer is that there is no cap, ask what the vendor's own inference bill looks like per store per month. A vendor who cannot answer that has not run the math on their own product.
Which tasks the agents took first
OpenAI classified its agent usage against a six phase taxonomy of AI research work published by Epoch AI: Decide, Design, Build, Run, Analyze, Communicate.
In January, research and infrastructure code dominated. By August, the notable increases were in technical help and monitoring runs. And this: "High-level planning still remains a minimal fraction of agent output tokens."
There is a lovely human detail underneath it. "Multiple teams which previously held office hours to help researchers troubleshoot their experiments have noted declining attendance in 2026, and one has stopped holding sessions entirely."
That is the shape of the whole thing. Agents ate the maintenance, monitoring and troubleshooting layer first. They did the work nobody enjoyed and everybody needed. Planning came last and still has not arrived.
What this means for your dealership
Autonomy claims should be priced by task horizon. An agent handling one inbound lead reply is doing a bounded task with a visible answer. An agent owning a 30 day follow up cadence across 400 leads, deciding who gets a call, deciding when to stop, deciding what to escalate, is doing the long horizon work where OpenAI's own intervention rate climbed. Vendors demo the first and sell the second.
The staffing math is multiplication. OpenAI did not report a smaller research organization. It reported the same people running more experiments, with August 2026 the all time high for experiments per active experimenter since tracking began in January 2025. If a vendor's pitch to you is headcount reduction, they are promising something the frontier lab with 3.1 agent workdays per human workday has not claimed for itself.
Point agents at your maintenance layer, because that is the layer where they already work. Inventory feed hygiene. VDP data completeness. Missing photos and thin descriptions. Ad account anomaly checks. Review response drafts. Broken redirect sweeps after a platform change. That is also the exact layer that decides whether AI systems can read your store at all, which is why it mattered when WebMCP shipped in August and why shopping agents fell over on retailer product feeds.
Measurement just became a buying criterion. OpenAI reports task success rate bucketed by how long the task would take a human, plus the share of successful tasks that needed intervention. Any vendor running agents in your store has those two numbers sitting in their logs. Asking for them is now a reasonable question with a public precedent behind it.
The counter-case, because it is a real one
The obvious objection: OpenAI's researchers are doing frontier AI research and your BDC is answering "is this one still available." Those are not equally hard tasks. An agent that needs steering on an 8 hour research task might need none on a 40 second lead reply. That is fair, and anyone who tells you OpenAI's intervention rate transfers straight onto your store is selling you the mirror image of the hype.
Two things survive the objection.
The shape of the curve transfers even when the number does not. Intervention rose with task length in OpenAI's data. Whatever your vendor's true rate is, it is higher for the long jobs than the short ones, and the long jobs are what gets sold as transformation.
And OpenAI published an unflattering number about its own product when nothing forced it to. When a vendor with a fraction of that capability publishes no number at all, the absence is itself information. We have written before about why dealers keep firing the wrong vendor, and the pattern is the same one: the thing nobody measured is the thing that was broken.
What to do about it
- Ask every AI vendor in your store for two numbers in writing: task success rate by task length, and the share of successful tasks that required a human to intervene.
- Write a task horizon into the pilot. Bounded tasks with a checkable output, for 60 days, before anything long running touches a customer.
- Get the usage cap in writing, and what the overage costs.
- Point the first pilot at maintenance work: feed accuracy, VDP completeness, redirect sweeps, review responses. Measure it against the same work done manually last quarter.
- Keep a named human on priorities and escalation. OpenAI kept theirs.
- Add the website and the data feeds to your next vendor review. An agent can only act on what your platform actually exposes.
Here is the useful takeaway. The company that builds the models your vendors resell now publishes a scoreboard for its own agents, including the parts that look bad. That scoreboard is the standard to hold your vendors to, and most of them are nowhere near it.
If you want to know how legible your inventory and your store actually are to the AI systems doing the reading, book a walkthrough and we will show you what they see.
Get Our Answers in Your Google Results
Add Savvy Dealer as a preferred source and Google highlights our articles with a Preferred badge in AI Overviews, AI Mode, and Top Stories. One click, then check the box next to savvydealer.com.
Ready to Transform Your Dealership's Marketing?
Schedule a free demo to see how Savvy Dealer can help you sell more cars.