Report 01 · Study materials
Methods and sources
What was compared, what was recorded, and what remains unknown.
Study setup
The experiment used five configured Claude model labels: Haiku 4.5, Sonnet 5, Opus 5, Fable 5, and Fable 5.1. Each model received a forecast email, a One-pager, and a slide deck assignment twice, producing 30 sessions. Matched assignments used identical fictional sales records and instructions. Each run began in a separate workspace.
The work was recorded on 3 September 2026. Cost figures, activity analysis, and displayed output observations use Run 1. The main time comparison includes both runs; the output and model pages provide both in full. Haiku remained cheapest in Run 2, while some other cost rankings changed. These are observations from a fixed data set, not expected savings or a permanent model ranking.
| Forecast email | An internal forecast call, supporting evidence, and management actions; fewer than 350 body words. |
|---|---|
| One-pager | A landscape page for a weekly revenue review, including forecast, movement, risks, and actions. |
| Slide deck | Six to eight landscape pages for revenue leaders, with source notes and missing information. |
| Source records | Current and prior sales records, targets, seller notes, and forecast rules. All companies and people in those records are fictional. |
Output review
On 5 September, a limited review checked six headline figures per file against the source calculations. Four values were incorrect across two outputs: three in the Haiku Run 1 deck and one in its Run 2 email. Related wording observations were recorded separately. These results do not validate every calculation or statement in a file.
The Haiku Run 1 email and both Fable 5.1 emails missed the word limit. Sonnet’s Run 2 email stopped at its spending limit after saving its unchanged file. Its reported cost and time end at that stop; later activity is excluded.
On 6 September, we inspected the Run 1 One-pagers and decks as rendered pages and read the original emails. The design comments describe visible layout and formatting. Reviewers could see model names. No blind review assessed overall quality, and no numerical design score was assigned.
Cost calculation
Four token charges are calculated from saved usage and prices: other input, input saved for reuse, reused input, and generated output. Their sum, plus active session time at $0.08 per hour, reproduces every recorded task cost when rounded to cents. The token prices were saved on 3 September; the active-time rate was verified from official documentation on 5 September and was not in the original price file.
Completion time measures the elapsed wait from the first task message to the recorded stop. Billed active time is a separate service measure. Tools have no separate recorded dollar charge or exclusive measured execution time. A request’s charge therefore cannot be divided among its tool calls from these records.
The preview calculation divides request token charges at the first recorded image-opening call. The later share includes requests with a higher request number and excludes session time. It includes all later activity, so it is not a measured price for revision alone.
The activity exhibit groups all requests from the five Run 1 One-pagers using reviewed tool instructions. Document work includes writing or changing PDF programs, rendering pages, reading document code, opening previews, and checking file properties or text. A command can combine several document actions; those charges remain together. Mixed document and non-document actions remain separate, as do responses without tools. Each amount includes the entire request’s input and output. It does not measure the tool’s own cost, time, or potential savings.
Token cost breakdown
Sonnet 5$1.81
- Other input157 tokens × $2.00 / million
- $0.0003
- Save input102,937 tokens × $2.50 / million
- $0.2573
- Reuse input5,067,826 tokens × $0.20 / million
- $1.0136
- Model output52,647 tokens × $10.00 / million
- $0.5265
- Active session time744.9 seconds × $0.08 / hour
- $0.0166
Total, rounded to cents $1.81
Fable 5.1$5.33
- Other input86 tokens × $10.00 / million
- $0.0009
- Save input100,669 tokens × $12.50 / million
- $1.2584
- Reuse input2,491,809 tokens × $0.25 / million
- $0.6230
- Model output68,554 tokens × $50.00 / million
- $3.4277
- Active session time1033.7 seconds × $0.08 / hour
- $0.0230
Total, rounded to cents $5.33
Token prices saved on 3 September; active-time rate verified on 5 September. Charges are shown to four decimals; the complete calculation is rounded once to cents. Source: recorded token counts, prices, and active seconds.
Evidence limits
Model names identify configured sessions. Individual request records do not identify the serving model, so calculations apply each session’s saved model price to its recorded usage. Complete request inputs were not preserved. Exact reuse of particular instructions or tool results is unknown. The main page’s request-loop and caching diagrams explain the process; they do not reconstruct a recorded request.
Two runs in a fixed order cannot establish typical variation, causes for every action, or future savings. A grouped command can contain several operations, and a response can request several tools. Neither grouping nor preview counts establishes simultaneous execution, effectiveness, or complete review.
The test software did not explicitly set the effort control, which influences token use. Records show high effort for four settings and no value for Haiku. Three early Run 2 sessions had spending limits that differed from the plan. Effects beyond the recorded Sonnet email stop were not measured. The downloaded procedure documents these setup differences.
Future research will test common assumptions about applied AI and quantify the effects of recommended practices. We will compare the quality, usage cost, and completion time of the same work under different setups. Skills, templates, independent checks, MCP and API access, tool-call efficiency, model choice, conversation history, and caching are planned comparisons. Their effects on quality, cost, time, and consistency remain to be measured.
Original materials
- Complete study packageOriginal assignments, data, outputs, and calculation scripts
- Procedure and setup differencesHow the original work was run
- Recorded model settingsModel labels, effort, order, and spending limits
- Numerical reviewFindings and limits of the six-figure check
- Editorial evidenceSources for the September 6 trace and design observations
- Activity groups and chargesWhole request charges grouped by work, for the five Run 1 One-pagers
- Reviewed activity labelsCall-by-call labels and source references for the grouping
- Complete request recordsTool inputs, returned results, and request usage
- Cost calculationsToken prices, request charges, and session time
The package’s offline checks need Python, not a model account. New experiments require compatible service and model access. The public records shorten long excerpts and replace private locations. Original outputs remain unchanged.
Both runs
Recorded task costs in US dollars. Review and stopping qualifications above apply to these values.
Forecast email
| Model | Run 1 | Run 2 |
|---|---|---|
| Haiku 4.5 | $0.09* | $0.12* |
| Sonnet 5 | $0.26 | $0.36* |
| Opus 5 | $0.46 | $0.48 |
| Fable 5 | $0.63 | $0.74 |
| Fable 5.1 | $1.09* | $1.29* |
* See the numerical, length, or spending-limit qualification above.
One-pager
| Model | Run 1 | Run 2 |
|---|---|---|
| Haiku 4.5 | $0.19 | $0.24 |
| Sonnet 5 | $1.81 | $1.56 |
| Opus 5 | $4.94 | $2.86 |
| Fable 5 | $2.59 | $3.32 |
| Fable 5.1 | $5.33 | $5.06 |
* See the numerical, length, or spending-limit qualification above.
Slide deck
| Model | Run 1 | Run 2 |
|---|---|---|
| Haiku 4.5 | $0.33* | $0.31 |
| Sonnet 5 | $3.78 | $2.85 |
| Opus 5 | $5.95 | $7.81 |
| Fable 5 | $7.67 | $5.57 |
| Fable 5.1 | $5.32 | $4.16 |
* See the numerical, length, or spending-limit qualification above.
Official references
- How Claude uses tools
- Prompt caching
- Token counting
- Claude reasoning
- Tenex
- MCP architecture
- Efficient MCP use
- Designing effective tools
- Loading tools when needed
- Model cost and performance
- Token and active-time pricing
Mechanism and company references checked on 6 September 2026. Experiment calculations retain the saved study prices.
Further research
Uber’s software-factory research breaks spending into model requests, token use, and prices. Its examples compare tool loading, repeated status checks, and returning smaller results. Ramp’s Inspect reports describe giving agents a working environment for tests and visual checks, then directing them at specific failures.
These company reports inform future comparisons. Their reported savings and performance do not measure the results of this study.
Company-authored engineering reports, accessed 6 September 2026.
Source inventory: 30 original output files, with hashes preserved in the study package.