Measuring agent workflows on public benchmarks
An agent that books invoices, answers questions over a database or fills in spreadsheets is more than a language model. It is a workflow: the model, the instructions it is given, the tools it may use and the steps it takes. Every part of it can be changed, and every change has a price. Measuring a workflow on a fixed set of cases is the statistical counterpart of a regression test, and it is what makes a change safe to deploy. This article shows how Dvergr, our open-source system for running and measuring agents, measures such workflows, what it found when we measured ten models on two public benchmarks, and why the most useful thing a measurement finds is often not the best model but a better workflow.
The question behind a leaderboard
Suppose a firm wants an agent to answer its staff’s questions about the company database. It has to choose a model, write the instructions, and decide what the agent may do: write one query and stop, or look at the data first, try a query, read the error and try again. A public leaderboard answers a narrower question, which model scored highest on someone else’s cases, with someone else’s instructions, on one run.
Dvergr answers the firm’s question directly. It takes a set of cases, each a task with the answer a person recorded for it, and a set of candidates, each a model together with its instructions and tools. It runs every candidate on every case and reports, for each candidate:
- how often it is right, with a range that says how much that number could move on another set of cases of the same kind;
- where it fails, check by check;
- what a correct answer costs at the model’s list price;
- how long it takes.
A firm that deploys an agent needs what software teams get from regression tests: a fixed set of cases with recorded answers, run again on every change. An agent’s output varies from run to run, and a provider can change the model behind a name it keeps, so one passing run says little. The result of a run is therefore a rate with an interval rather than pass or fail, and a change is judged by comparing it with the previous version on the cases where the two disagree.
We tried this on two public benchmarks, where anyone can check the cases and the answers. SpreadsheetBench asks for changes to Excel workbooks. BIRD asks questions in plain language about real databases, to be answered with a query.
The models come in families, named here from cheapest to most expensive: OpenAI’s Luna, Sol and Astra, in generations 5.6 and 6 (6.1 for Sol); Anthropic’s Claude Haiku, Sonnet and Opus 5.5; and two models with published weights, GLM and DeepSeek.
Results at a glance
Each dot is one model. Its bar shows how far the score could move if we had picked a different 100 questions of the same kind (59 tasks on SpreadsheetBench), the way a poll of 100 people would come out differently with another 100. Each model answered each question once, and the bar follows from that count alone, which is why it is nearly the same length for every model: with 100 questions and scores around two thirds, a score can move by about nine points either way. Where two bars overlap, the measurement cannot tell the two models apart, and the choice between them comes down to cost and time, on the horizontal axis.
Because every model answered the same questions, two models can also be compared more sharply. Most questions do not help, because both models got them right or both got them wrong. What decides is the questions where they differ. If model A was right and B wrong on twelve questions, and the reverse happened on only two, A is better even though their bars overlap; six against five says nothing. We give these counts with a p-value, the chance that two equally good models would split their differences at least this unevenly, and below 0.05 we treat a difference as real. Dvergr’s report gives the bars; the counts and p-values here we computed from the per-question results it records.
The tables also give the median time per case. Time depends on how busy the provider was that day, so it compares candidates run together better than candidates run on different days.
How a measurement runs
A benchmark enters Dvergr as a case pack: the cases, their recorded answers, and a checker that decides whether an answer is right. Before any model runs, the pack is certified. Certification lists every case that cannot grade an answer and says why: an id that is missing or appears twice, an empty answer, the same inputs recorded with different answers, or a recorded answer the checker itself does not accept. Those cases are set aside. On SpreadsheetBench, certification kept 390 of the 400 tasks without spending any model tokens.
The measurement itself takes place in a room, Dvergr’s unit of shared state: its files, its database and its history. For each candidate and case, Dvergr makes a fork of the room, a private copy that costs almost nothing to make because it shares everything until something changes. The candidate works in its fork: the files and the workbook it changes belong to that copy, and a database it may only read is reached for each query through a fresh read-only connection or an unchangeable snapshot, so nothing it does reaches the original data or the candidates working beside it. The checker grades the attempt’s result (the files in the fork, the query it submitted, or its edits replayed on a fresh copy of the workbook), the verdict is recorded in the room, and the fork is thrown away. Dvergr’s own share of each case, the fork included, is a fraction of a second, against the ten to seventy seconds a model takes to answer.
The statistics depend on these forks. Because every candidate starts from the same state, two candidates’ answers to a case can be compared directly, which is what the paired comparison needs. Because no attempt can change what another sees, a candidate is graded on its own work. And because a fork shares everything that has not changed, running every candidate on every case stays cheap. This follows from how Dvergr keeps state: the room’s files and its database are forked with copy-on-write branches across git and Datahike, so a fork records only what it changes and never writes to the original.
Four rules keep the numbers honest:
- The same cases for everyone. Every candidate answers every case, which is what makes the paired comparison possible.
- A broken connection is not a wrong answer. When the provider is down or a request times out, the case is run again. A model that answers wrongly, or gives no answer, gets its verdict.
- List price. Models are billed per token, a piece of a word, read or written. Cost is what the same tokens would cost through the provider’s public price list, whether a subscription or the firm’s own key paid for them, with text the provider has already seen recently (cached input) at its lower rate.
- A changed setup is a new experiment. The cases, candidates, instructions, checker and software versions together define an experiment. An interrupted run resumes where it stopped, but if any of them has changed, running it again starts a new experiment instead of mixing two.
Does the agent need to explore?
The candidates above are models inside Dvergr’s own workflows. On BIRD, the model may run queries in SQL, the standard database query language, see their results or errors, and submit an answer when it is satisfied, within twenty turns. On spreadsheets it may read cells, write values and formulas, see what they compute, and submit.
A firm should know how much of a score is the model and how much is the workflow around it. An agent that explores uses more tokens than one that answers at once, and it is worth knowing what those tokens buy. So we also ran each benchmark’s own reference setup, with the same model on the same cases. These runs are not Dvergr workflows: they use each benchmark’s published code to build the prompts and grade the answers, and Dvergr only connects them to the model, so that both setups reach it the same way. The scripts, with the steps to reproduce each number below, are in the repository under benchmarks/reference/.
On BIRD, one instruction did most of the work
BIRD’s reference setup is a single prompt: the list of the database’s tables and columns (its schema), the question and a hint, and a request to write the SQL query. The model answers once, without seeing any data. We ran it as published, then with one line added, return exactly the columns the question asks for. Graded with BIRD’s own evaluation script, on the same 100 questions:
| Workflow | GPT-5.6 Luna | Claude Haiku 5.5 | Tokens per question (Luna) |
|---|---|---|---|
| BIRD’s prompt, one reply | 51 | 53 | 1,200 |
| the same, plus the line about columns | 61 | 64 | 1,200 |
| Dvergr: run queries, see results and errors, then submit | 65 | 69 | 8,200 |
Two questions show what each part contributes. Question 781 asks for the heights of the heroes whose eye colours are amber. With BIRD’s prompt, Luna returned each hero’s name beside the height. The heights were right, but BIRD compares the returned table as a whole, and a table with an extra column is wrong. With the line about columns it returned the heights alone and passed. Of the fifteen questions Luna answered on our workflow and missed with BIRD’s prompt, seven failed for no other reason than such an extra column, and the one line is worth ten points for both models, a difference too large for chance (p = 0.013 and 0.003).
Question 817 asks for the race of the blue-haired male superhero, and its hint says the colour is written 'blue' and the gender 'male'. In the database they are written 'Blue' and 'Male', and SQLite compares text exactly. The one-reply query followed the hint and found nothing. In Dvergr’s workflow Luna wrote a query that ignores capitals, ran it, saw three heroes, and submitted the query for their races. That is the kind of question where looking at the data pays.
Exploring the data adds another four or five points. With 100 questions that could be chance (p = 0.34 and 0.30), and it costs seven times the tokens. For these questions the cheapest good workflow is a single reply with the right instruction, at about a third of the cost per correct answer. The agent that explores scores higher, but 100 questions cannot show that the gain is real, and it costs more.
We had written that line into our own workflow from the start, because we knew how BIRD grades. Until we measured it, we did not know it was most of what our workflow added.
On spreadsheets, the recalculating program decided the score
On SpreadsheetBench the scores depended less on the model than on how the answers were checked.
The benchmark’s reference setup shows the model the instruction and the first five rows of each sheet, and asks for a short Python program that edits the workbook. Its multi-round version runs the program and shows the model the output or error, for up to five turns. We ran both versions with GPT-5.6 Luna on our 59 tasks, each program in a Docker container that can see only its own input file, and graded them with the benchmark’s own comparison code. Dvergr’s workflow writes no Python: its agent edits the workbook directly, through a spreadsheet engine in Dvergr’s sandbox.
SpreadsheetBench’s grader reads each answer cell’s stored value, the result a spreadsheet program saved with the formula, and never computes a formula itself. A file written from Python stores formulas without values, so the benchmark’s instructions add a step before grading: open every workbook in LibreOffice (on Linux and macOS) or Excel (on Windows) and let it recalculate. We followed the instructions, using LibreOffice on Linux.
One task shows what happened. Task 56563 has a column of monthly amounts, one of which is an error value, and asks for a formula that adds them up anyway. Both workflows wrote the same correct formula into the answer cell, =AGGREGATE(9,6,C2:C13), which sums the range and skips cells holding errors. Its value is 77,772.6. We graded that one cell two ways:
| Graded by | What it found in the cell | Verdict |
|---|---|---|
| the grader after LibreOffice recalculated, as the benchmark’s instructions say | #NAME?: LibreOffice recognises functions added to Excel since 2010, such as AGGREGATE, only when the file marks them with a prefix (_xlfn.AGGREGATE), and neither workflow wrote it | wrong |
| Dvergr’s spreadsheet engine | 77,772.6, as does LibreOffice once the prefix is written | right |
The same answer was wrong or right, depending on which program computed it. Across the 59 tasks:
| Workflow | Recalculated by LibreOffice (the benchmark’s instructions) | Tokens per task (median) | Median time |
|---|---|---|---|
| SpreadsheetBench, one reply with code | 34 | 2,200 | 15 s |
| SpreadsheetBench, up to five turns with execution | 41 | 4,300 | 17 s |
| Dvergr: read and write cells, see what they compute, submit | 41 as written, 48 as Excel stores them | 5,700 | 24 s |
Graded the benchmark’s way, after LibreOffice’s recalculation, the multi-round workflow scores 41, and Dvergr’s workbooks as written also score 41; each is right where the other is wrong on eight tasks, so on that footing they tie. Written with Excel’s prefix, Dvergr’s workbooks score 48. Graded by Rechentafel, the open-source spreadsheet engine Dvergr’s agent works in, they score 51. Grading with our own engine would prove little if the engine had not been checked against Excel. The benchmark’s recorded answers are the values Excel saved, and for 390 of the 400 tasks Rechentafel recalculates each recorded answer to exactly that value, which is how certification works on this benchmark. The workflows differ less than the graders do. A firm measuring a spreadsheet workflow should grade with the program its people use, or with one checked against it.
Improving a workflow by measuring it
The line about columns is the kind of change a measurement should find, not one an author should have to know in advance. Dvergr is built so that variants of a workflow can be written and measured like any other candidate.
A workflow in Dvergr lives in its room as ordinary files: the task and its instructions, the checker, the cases, and any code the agent runs in its sandbox. Because they are files in the room, an agent working there can read them and write a variant: a sentence added to the instructions, a look at the data before the answer, a different model. Each variant is a candidate like any other, measured in its own forks on a set of tuning cases. A variant that does better is then measured on cases it has not seen, so that one tuned to the quirks of one set of questions does not pass for an improvement. A variant’s checker is trusted only after the host’s administrator promotes it, through an administrator’s connection. The promotion is recorded outside the room’s files, where no agent can write, so an agent cannot raise the trust of its own checker.
We took one step of this loop by hand on our Datalog workflow for BIRD. Datalog is the query language of Datahike, the database Dvergr keeps its own state in, and the workflow answers BIRD’s questions with it instead of SQL. After reading the failures of one run we added two sentences to its description of the language, and measured the change on another 100 questions that run had not seen:
| Workflow | Correct (of 100) | Input tokens per question |
|---|---|---|
| SQL on SQLite | 59 | 8,880 |
| Datalog, revised description | 63 | 14,042 |
| Datalog, previous description | 58 | 16,684 |
The revision answers five more questions with less input, but at 100 questions that could be chance (p = 0.23). That is the honest result of one step of tuning: a promising change, to be confirmed on more cases before it is adopted. Today a person or an agent proposes each variant and Dvergr measures it. Letting an agent in the room run the whole loop, propose, measure and hand the winner to the owner, is what we are building next. The rule for adopting a variant is the one a team applies to a code change: merge it when the cases show it is not worse.
SpreadsheetBench in detail
This section and the next give the full results behind the chart, model by model. These numbers use our workflow and our spreadsheet engine, on 59 of the 390 certified tasks, so they are not comparable to the public leaderboard, where agents work with other tools.
| Model | Tasks | Passed | 95 % interval | Cost per pass | Median time |
|---|---|---|---|---|---|
| GPT-5.6 Luna | 59 | 51 | 76–93 % | $0.0026 | 24 s |
| GPT-5.6 Sol | 59 | 52 | 78–95 % | $0.033 | 27 s |
| Claude Opus 5.5 (Claude Code subscription) | 59 | 52 | 78–95 % | $0.062 | 17 s |
| Claude Haiku 5.5 (Claude Code subscription) | 59 | 47 | 68–88 % | $0.0046 | 22 s |
| Claude Sonnet 5.5 (Claude Code subscription) | 59 | 46 | 66–87 % | $0.33 | 19 s |
| GPT-6 Luna | 59 | 48 | 70–90 % | $0.0008 | 10 s |
| GPT-6.1 Sol | 59 | 55 | 85–98 % | $0.010 | 9 s |
| GPT-6 Astra | 59 | 55 | 85–98 % | $0.050 | 9 s |
Paired with GPT-5.6 Luna, none of the GPT-5.6 and Claude models is measurably different. Opus is right where Luna is wrong six times and the reverse five times; Sol three and two; Haiku four and eight; Sonnet three and nine (the closest, p = 0.15). What differs is the cost. Per passing task Haiku costs about twice as much as Luna, Sol thirteen times, Opus twenty-four times and Sonnet over a hundred times, because Sonnet sometimes wrote very long answers: its slowest tenth of tasks took over nine minutes.
The GPT-6 models, run on the same tasks a few days later, are the first to move the top. GPT-6.1 Sol and GPT-6 Astra pass 55 each and disagree on only four tasks, two each way, so Astra costs five times as much as Sol for the same answers. Sol 6.1 is measurably better than GPT-6 Luna (eight tasks to one, p = 0.039) but not than GPT-5.6 Luna (five to one, p = 0.22). They were also about three times faster, partly because of the models and partly because of the day.
Every task but one was solved by at least one model, so most failures belong to a model, not to the task.
BIRD in detail
BIRD’s questions cover 11 databases. We can answer each question three ways: SQL on SQLite, the database BIRD ships with; SQL on Datahike, through its PostgreSQL-compatible layer pg-datahike; and Datalog on Datahike. Even before any model runs, 1372 of BIRD’s 1534 reference queries return exactly SQLite’s rows on pg-datahike.
With GPT-5.6 Luna on 100 fresh questions held out from tuning:
| Engine | Correct (of 100) | 95 % interval | Cost per correct answer | Input tokens per question | Median / slowest tenth |
|---|---|---|---|---|---|
| SQL on SQLite | 65 | 55–74 % | $0.0023 | 7,839 | 10 s / 23 s |
| SQL on Datahike (pg-datahike) | 65 | 55–74 % | $0.0026 | 10,136 | 10 s / 41 s |
| Datalog on Datahike | 64 | 54–73 % | $0.0044 | 15,061 | 13 s / 42 s |
The three engines answer equally well: each pair disagrees on seven to ten questions, split almost evenly. They differ in cost. In Datalog the model looks at the schema more before it answers, so a correct answer costs about twice as much as in SQL. Which query language to give an agent is a cost decision, and the measurement says how large it is.
Seven models on the same 100 questions:
| Model | SQL on SQLite (of 100) | Datalog (of 100) | Cost per correct answer, SQL / Datalog | Median time, SQL |
|---|---|---|---|---|
| GPT-5.6 Luna | 65 | 64 | $0.0023 / $0.0044 | 10 s |
| Claude Haiku 5.5 | 69 | 68 | $0.0026 / $0.0047 | 9 s |
| DeepSeek V4.1 Flash (open weights) | 66 | 64 | $0.0074 / $0.0195 | 8 s |
| GLM 5.3 Flash (open weights) | 64 | 58 | $0.0042 / $0.0106 | 6 s |
| GPT-6 Luna | 68 | 60 | $0.0009 / $0.0014 | 5 s |
| GPT-6.1 Sol | 74 | 71 | $0.015 / $0.019 | 8 s |
| GPT-6 Astra | 75 | 70 | $0.067 / $0.091 | 6 s |
Among the first four models, no two are measurably different. Any two disagree on 10 to 18 questions; the most lopsided pair (GLM’s Datalog against Haiku’s, p = 0.013) is no more lopsided than chance alone produces somewhere among 28 comparisons.
GPT-6.1 Sol and GPT-6 Astra answer 74 and 75 in SQL, against 65 for GPT-5.6 Luna (eleven questions to two and twelve to two), and tie with each other. Counting all the comparisons we made, even that is not quite conclusive, but it is the first gain we have seen, and it comes within about five questions of the 80 that BIRD’s recorded answers allow (see below). Again Astra costs about five times as much as Sol per correct answer.
The two open-weight models (models anyone may download and run on their own machines), served by Fireworks, answer as many questions as GPT-5.6 Luna and Claude Haiku, at two to three times the cost per correct answer, because they used more tokens and none of their input was served from a cache. For a firm whose data may not leave its own machines, that is the useful result: a model it can run itself answers these questions as well, and its cost is then the firm’s hardware, not a list price.
What the benchmarks get wrong
A benchmark is only as good as its recorded answers and its grader, and both benchmarks have faults that change scores.
On BIRD, twenty of the 100 questions were answered correctly by none of the first four models in either language. On most of them the models agree with each other and not with the recorded answer:
| Question (BIRD id) | Recorded answer | What the models answered |
|---|---|---|
| atoms of molecule TR346 and its bond types (309) | a query for molecule TR000 | molecule TR346 |
| the German type of a card (482) | the English type | the German type |
| patients with abnormal CRP and no recorded data (1256) | 208, the number of lab rows | 25 patients; the only answer graded right, GPT-6 Luna’s in Datalog, repeated the mistake |
| expenses of the budget with the lowest remaining (1365) | one row, cut off by LIMIT 1 | both expenses of that budget |
| difference of two percentages (1458) | 12.12 | 0.1212, following the hint’s formula, which has no ×100 |
| eye colours of Marvel heroes by popularity (728) | colour, count and a rank column | colour and count |
We count about ten questions whose recorded answer is wrong, four more whose hint contradicts it, and the rest a matter of which columns or which wording. On these 100 questions the best attainable score is about 80, not 100. Paired comparisons survive this, since every candidate loses the same questions, but the differences shrink, and a model that answers the question as asked gets no credit for it. Our grader agrees with BIRD’s own on all 200 SQL answers we checked, so the scores above are the benchmark’s own.
On SpreadsheetBench:
- The published evaluation script cannot score the published tasks. On the repository’s main branch it compares each task’s input workbook with the recorded answer (the line that reads the model’s output is commented out), and its file names are those of an earlier release. Every published score on the 400 verified tasks comes from a modified script that is not part of the benchmark.
- Formulas are graded as empty unless the workbook is recalculated first, as described above. The benchmark’s instructions include that step, but the program that recalculates changes the score, and the step also recalculates the recorded answers, so after it they are that program’s results too.
- One task can never pass. The answer position of task 45944 contains spaces, and the checker fails on it whatever the answer.
- Some recorded answers are wrong or misfiled. Task 118-50’s answer misses a pair its own rule finds in the input, and the two strongest models were graded wrong for finding it. Task 42930’s answer file carries another task’s number.
Certification sets aside ten of the 400 tasks for reasons like these before any model runs.
Checking our own measurements
The setup that measures is code too, and it needs the same tests as the workflows it measures. Before publishing, four AI review agents read our benchmark code, documentation and stored results, each with one question: whether every benchmark runs the same way, whether the forks isolate each attempt, whether the documentation matches the code, and whether the numbers are right. They found nine problems, among them a provider outage counted as the model’s failure and cached input billed at its full price. Each fix came with a test that fails without it, and the affected runs were repeated or withdrawn.
What these numbers cannot tell you
Each model answered each case once, so a few answers could change on another run; the intervals say how much that matters. The numbers describe these cases with these workflows, and a different prompt, tool or model version can change them, which is why an experiment records all three. A tie at 59 tasks says what 59 tasks can separate, not that two models are equal. As with tests, coverage sets what can be found: with 100 cases, a regression of a few points can go unnoticed. And a public benchmark is not your workflow: it shows that the method works where others can check it, not how a model will do on your cases.
Where the method comes from
Seen this way, a workflow is a probabilistic program: a program whose output is drawn from a distribution, and measuring it is inference about how often that output is right. The Jeffreys interval in our reports is a Bayesian estimate of that rate. The view comes from probabilistic programming, which Christian Weilbach worked on in his doctoral research in Frank Wood’s group, with the Anglican language and the Daphne compiler (TMLR 2025). It continues in Foerster, our library for inference over forkable worlds, where alternatives are forks of the same state, as candidates are here. Dvergr’s measurements do not use Foerster yet. Two steps it would allow are stopping a comparison as soon as the evidence decides it and spending more cases where two candidates are close.
Try it on your own cases
The benchmark that matters for a firm is its own history: past bookings, coded invoices or resolved tickets, each with the outcome a person recorded. Dvergr is open source and runs on your own machine:
- Install and connect. Clone the repository (it needs the Clojure CLI and babashka), and add
bin/dvergr-mcp --profile benchas an MCP server in Claude Code or Codex. The models it compares can be used through a Claude Code or Codex subscription or any API key; the README lists the providers, the setup and where Dvergr keeps its state. - Turn your history into a case pack. Export a table of past cases as CSV, as Excel or DATEV write it, and name its columns: the id, the inputs, and the outcome with a rule for checking it (exact, a number within a tolerance, a date, a set). Certification reports which cases can grade an answer and why the others cannot, which is a first result in itself.
- Run the comparison. Ask your assistant to run
catalog_benchmarkwith the models you want compared, or, from a shell, turn the CSV into a pack withclojure -M -m dvergr.catalog.casepack-cliand runclojure -M -m dvergr.catalog.room-run <pack> --models … --cases 30. The report gives each model’s success with its interval, where it fails, the cost per correct result and the time.
Running benchmarks and case packs in the documentation give the details. A following article applies this to a year of bank bookings exported from DATEV.
If you would rather we built the benchmark with you, we run fixed-fee pilots: six weeks on one workflow of yours, ending with the benchmark and a report on which model, instructions and steps to run, at what cost. Send us a note.
Method
| Models | GPT-5.6 Luna and Sol, GPT-6 Luna, GPT-6.1 Sol and GPT-6 Astra (Codex subscription); Claude Haiku 5.5, Sonnet 5.5 and Opus 5.5 (Claude Code subscription, each pinned to its exact version); GLM 5.3 Flash and DeepSeek V4.1 Flash (Fireworks). Each answers every case once. |
| Cases | BIRD dev, the held-out third of each database’s questions, fresh samples of 100; SpreadsheetBench verified, held-out tasks among the 390 certified |
| Grading | BIRD: the returned rows equal the recorded query’s rows; SpreadsheetBench: the answer cells equal the recorded workbook’s values |
| Statistics | Jeffreys 95 % intervals for success rates, as Dvergr’s report gives them; exact McNemar test on the cases two candidates disagree on, computed from the recorded per-case verdicts |
| Cost | list price per million tokens on the day of measurement, cached input at the cache rate |
| Reference setups | BIRD’s prompt from its published gpt_request.py and its evaluation script; SpreadsheetBench’s single- and multi-round drivers and its comparison code, the generated programs run in a Docker container; both reach the model through Dvergr’s model connection: benchmarks/reference/ |
| Code and data | BIRD and SpreadsheetBench are public; Dvergr and both benchmark adapters are open source: github.com/replikativ/dvergr |
See how Simmis lets teams delegate consequential work without losing control of what becomes official.
simmis