Benchmarking your own agent spend
DoorDash's model spend rose about 20× in five months, so they built a benchmark over their own coding tasks to work out what it bought — and found that the models crush a scrubbed version of a task and then underperform on the real data. The first buyer-side view in this notebook, and the open question it leaves.
From the NoPriors interview. Andy Fang on what a 10,000-person company does after a year of letting people spend on models freely. It is the first source here written from the buyer's chair rather than the vendor's, which changes which questions get asked.
The numbers, such as they are
June's spend was about 20× January's. Asked whether it is still climbing, he says it has flatlined — and attributes that to deliberate effort rather than saturation: the easy waste removed first, then the harder question of what the rest is returning.
Dashbench is their published benchmark, announced a couple of weeks before the recording. Two things about it are worth noting. It scores models and harnesses together on coding tasks, not models alone — which is the correct unit, and the one this whole cluster argues for. And its stated purpose is not leaderboard position but working out the ROI on the spend, including the specific move of routing cheap tasks to open-weight models and paying frontier prices only where they buy something. His phrasing for the target is frontier-level intelligence at less than frontier cost.
Also worth recording, because it is the part nobody publishes: seat growth is fastest in the non-technical organisations — analysts, operators, account managers automating merchant QBRs — while spend remains overwhelmingly engineering. They have no benchmark for that work yet. That is where they say they are going next, and it is the harder problem: coding tasks come with an exit code, and a QBR does not. Which is the whole difficulty, stated as an org chart.
The finding: scrubbed tasks pass, real ones don't
The part I would keep. Their loop with the frontier labs on non-coding work — accounting, analytics — runs into a consistent pattern:
- Ask the team how the models do on their task. Answer: "yeah, it works okay."
- Scrub the data, hand it to a lab, build an RL environment around the task. The models crush it.
- Point the same thing at real enterprise data with everything attached. It does not hold.
Fang's open question is whether that is a harness gap — things they have to build for the model to perform — or something the models genuinely lack in their data distribution. He does not resolve it, and the honest answer is that the interview contains no way to tell.
An answer to that question arrived on 2026-08-13, from someone with no stake in DoorDash's procurement. Greenblatt is asked the general form of it — does training on environments that don't resemble the real job produce real competence — and says the RL distribution already deviates a long way from how models get used in practice, and transfers anyway. On his account the answer to Fang's either/or is neither: not a missing harness and not a permanent hole in the data distribution, but a gap that closes as environments start being built out of production data rather than out of scrubbed samples. That is a prediction rather than a finding, and it is checkable — it says step 3 stops failing once step 2 is fed the messy version. Worth noting which way it cuts for a buyer: if he is right, the thing that fixes the models' performance on your real data is your real data.
Two readings, and they are not exclusive:
As evidence about the models, it is the enterprise instance of jaggedness observed from the outside. Step 2 is the verifiability operation — scrubbing and building an RL environment is exactly the work of turning a soft task into a gradable one — and the result is the model performing inside the rails and losing it outside. The gap between step 2 and step 3 is a measurement of how much of the real task the rails did not cover.
As evidence about the benchmark, it is worse news than it sounds, and it applies to Dashbench itself. If the models are graded on the scrubbed task, the benchmark reports on the scrubbed task. That is the closing worry on the jaggedness page — a benchmark suite is a map of what was optimised — arriving from the other direction: not "the gaps are unmeasured" but "the thing you built to measure ROI shares the blind spot you built it to find". The correction is unglamorous and expensive: grade on the messy data, and accept a benchmark that is harder to run and less comparable to anyone else's.
And the either/or itself is an old question. Added 2026-08-14: "harness gap or something the models lack" is the 2009 argument about whether the data or the architecture deserves the credit, asked about agents instead of about vision — and that argument was settled, when it was settled, by holding one side fixed and running the cheap baseline. Efros did it with nearest neighbours; the episode is here, along with the observation that nobody in the current round has run the equivalent. Fang's pattern is also dataset bias with the sign flipped — the benchmark forgiving because it is the version of the job a lab could grade rather than because the model was fitted to it — which is the same failure vision diagnosed in 2011 and did not fix.
Note also what step 1 says on its own. "Yeah, it works okay" is not a measurement, and the whole apparatus exists because self-reported satisfaction from the team using the tool turned out to be uncorrelated with anything. That is the same reason the maker cannot be the judge — one level up, with a department in the generator's seat.
Why this matters for a one-person version
Dashbench is out of reach for one person and the shape of the question is not. The notebook's evaluation thread has so far been about grading an artifact an agent produced (Generator–evaluator loops, Building a generator–evaluator harness: A practical implementation recipe). This is the other question, and it has never been asked here: what is the spend buying, and which model should do which job? The owner's version of that has a concrete form already — the desktop under the desk exists on precisely the bet DoorDash is now trying to price, that ordinary work runs locally on open weights and only the hard jobs go out to an API. See Speciation, and why we only ever touch the context window, where the same routing decision is discussed from the model side.
What transfers is the sequence rather than the tooling: you cannot route by cost until you have
a task set you actually care about, and you cannot claim a return until something on the output
side is counted. This repo has the first half — npm run build is the test, /lint is the
health check, and the sessions in log.md are a task history nobody has ever scored.
Where it thins out
A flatlining cost curve is not a return. No output measure is offered against the spend, so "we need to start calculating the ROI" is where this account stops, not a result it reports. The 20× is also unnormalised — over a period in which headcount using the tools, and the price per token, both moved.
Dashbench was named, not described — until two days before this interview, when the team that built it published the methodology. See the write-up, which answers most of what is missing here and corrects one thing: DashBench measures a code reviewer on replayed historical PRs, not coding tasks in general. Benchmarking coding agents is the next thing they want, not the thing they have. "The correct unit" survives the check — the harness is explicitly part of what varies, and cost and latency are recorded as results rather than controlled away.
It also partly answers the worry above about grading on scrubbed data. DashBench replays real PRs from the real codebase, with benign and later-reverted cases deliberately included, and its labels come from three disagreeing sources rather than one convenient one. That is what "grade on the messy data, and accept a benchmark that is harder to run" looks like when someone actually pays for it. What it does not do is close the gap Fang describes on the non-coding work — that is still where the scrubbed-versus-real problem lives, and still unmeasured.
The token price is not a property of the model either. One thing this page treats as given — that spend is spend, and the question is only what it bought — gets complicated by the provider-side view. Two of its findings bear directly on a buyer's arithmetic. Serving the same weights, the spread between a careless deployment and a careful one on identical hardware is 2–4×, so the provider spreads visible on OpenRouter or Artificial Analysis are a real quality difference and not only a pricing one. And past some volume the per-token model stops being the cheaper unit at all: at millions of tokens per hour, renting the box by the hour and taking responsibility for saturating it is the cheaper arrangement, with the trade being that saturation becomes your problem.
Neither changes the argument on this page, and both sharpen what a benchmark over your own tasks should be varying. If cost and latency are results rather than controlled-away parameters — which is the thing DashBench gets right — then the provider and the serving configuration are as much a part of the measured system as the model name, and a 20× spend curve has at least one term in it that is procurement rather than usage.
And the acquisition is the same claim in a different register. They bought a company (transcribed as Metis) to import AI-native habits, on the diagnosis that a large company's problem is not access but that people cannot imagine what is possible. Plausible, and also the kind of thing that is said after an acquisition. There is no before-and-after here, which is the one thing that would make it evidence.
A cleaner example of the same thinning is a vendor's own "under $6 an hour" figure for running a frontier model as an AI Engineer — a unit price (tokens/second × $/token), not a cost tied to a task set anyone could rerun. It skips past exactly the step this page is built around: a flatlining spend curve, or a cheap hourly rate, says nothing until something on the output side is counted against it.
Linked from
- Agentic engineering: the work moves to the harnessThe emerging discipline around long-running coding agents — designing the scaffolding, feedback loops and environments that let an agent do reliable work, rather than writing the code yourself. Entry point for the harness-design cluster in this notebook.
- Data or architecture: the control experiment nobody runsBefore AlexNet the field's best argument against neural networks was that the data was doing the lifting — and Efros had the control experiment to prove it. AlexNet settled that case on the same data everyone else had. The reason to keep the episode is that the modern version of the argument is running now, with the control missing.
- DoorDash on agentic commerce and DOT (NoPriors)Notes on the NoPriors interview with DoorDash's two co-founders — a natural-language ordering surface that half the time sends people to a merchant they have never used, a delivery robot built to a use case rather than to a capability, and an internal benchmark built to work out what a 20× rise in model spend actually bought.
- GPT-6 Astra: an automated AI Engineer for under $6 an hourA Latent Space writeup of early access to OpenAI's GPT-6 Astra, framed around the claim that it is cheap enough and capable enough to function as a junior AI Engineer — running 20-50 subagents in parallel, monitoring its own multi-day jobs, and building its own benchmarks. Read as a vendor demo rather than a measurement, and cross-checked against what the notebook already has on spend, throughput and unattended runs.
- Inference engineering as a discipline (Latent Space × Baseten, 2026-08)Philip Kiely and Ali Taha on what actually happens between an open-weight checkpoint and a production endpoint — the four optimisations that stack to 2–4× on fixed hardware, why quantising more of a model can make it better, and the reliability failures that turn out to be a race condition in a kernel rather than anything about the weights.
- Jaggedness: what RL optimises, and what stallsKarpathy's argument that models are advancing only where a reward can be computed, with a joke that hasn't changed in three or four years as the demonstration — and why that is the same boundary the harness cluster keeps running into from the other side.
- Learning from deploymentProduction traffic is becoming training data — tasks the model did badly on, turned into environments that match them exactly, with the rubric built from what the human actually wanted. Which is the loop this notebook already admires, run with the model as the artefact, and it changes where the reinforcement comes from.
- Optimising for the benchmarkComputer vision spent the 2000s carving its best models around the quirks of the datasets it was scored on, and hard negative mining was how it did it — training against your own error cases, one gradient step at a time. The episode is the oldest worked example in this notebook of a measure being optimised until it stops measuring.
- Speciation, and why we only ever touch the context windowKarpathy expects specialised models and does not see them arriving, and his explanation is that adjusting weights without losing capability is still an underdeveloped science while context windows just work. Plus where he puts the open-weight models — six to eight months behind, and better off there.
- The case for recursive self-improvement (Dwarkesh × Greenblatt, 2026-08)Ryan Greenblatt argues that AI research is verifiable enough to automate itself, that automating it buys four or five years of progress in one, and that what comes out the other end can transform the world without ever learning to play politics. His median for full automation is 2030–31 and his odds on takeover by 2040 are 35–40%.