Report 01 · Study materials

Methods and sources

What was compared, what was recorded, and what remains unknown.

Study setup

The experiment used five configured Claude model labels: Haiku 4.5, Sonnet 5, Opus 5, Fable 5, and Fable 5.1. Each model received a forecast email, a One-pager, and a slide deck assignment twice, producing 30 sessions. Matched assignments used identical fictional sales records and instructions. Each run began in a separate workspace.

The work was recorded on 3 September 2026. Cost figures, activity analysis, and displayed output observations use Run 1. The main time comparison includes both runs; the output and model pages provide both in full. Haiku remained cheapest in Run 2, while some other cost rankings changed. These are observations from a fixed data set, not expected savings or a permanent model ranking.

Forecast emailAn internal forecast call, supporting evidence, and management actions; fewer than 350 body words.
One-pagerA landscape page for a weekly revenue review, including forecast, movement, risks, and actions.
Slide deckSix to eight landscape pages for revenue leaders, with source notes and missing information.
Source recordsCurrent and prior sales records, targets, seller notes, and forecast rules. All companies and people in those records are fictional.

Output review

On 5 September, a limited review checked six headline figures per file against the source calculations. Four values were incorrect across two outputs: three in the Haiku Run 1 deck and one in its Run 2 email. Related wording observations were recorded separately. These results do not validate every calculation or statement in a file.

The Haiku Run 1 email and both Fable 5.1 emails missed the word limit. Sonnet’s Run 2 email stopped at its spending limit after saving its unchanged file. Its reported cost and time end at that stop; later activity is excluded.

On 6 September, we inspected the Run 1 One-pagers and decks as rendered pages and read the original emails. The design comments describe visible layout and formatting. Reviewers could see model names. No blind review assessed overall quality, and no numerical design score was assigned.

Cost calculation

Four token charges are calculated from saved usage and prices: other input, input saved for reuse, reused input, and generated output. Their sum, plus active session time at $0.08 per hour, reproduces every recorded task cost when rounded to cents. The token prices were saved on 3 September; the active-time rate was verified from official documentation on 5 September and was not in the original price file.

Completion time measures the elapsed wait from the first task message to the recorded stop. Billed active time is a separate service measure. Tools have no separate recorded dollar charge or exclusive measured execution time. A request’s charge therefore cannot be divided among its tool calls from these records.

The preview calculation divides request token charges at the first recorded image-opening call. The later share includes requests with a higher request number and excludes session time. It includes all later activity, so it is not a measured price for revision alone.

The activity exhibit groups all requests from the five Run 1 One-pagers using reviewed tool instructions. Document work includes writing or changing PDF programs, rendering pages, reading document code, opening previews, and checking file properties or text. A command can combine several document actions; those charges remain together. Mixed document and non-document actions remain separate, as do responses without tools. Each amount includes the entire request’s input and output. It does not measure the tool’s own cost, time, or potential savings.

One-pager · Run 1

Token cost breakdown

Sonnet 5$1.81

Other input157 tokens × $2.00 / million
$0.0003
Save input102,937 tokens × $2.50 / million
$0.2573
Reuse input5,067,826 tokens × $0.20 / million
$1.0136
Model output52,647 tokens × $10.00 / million
$0.5265
Active session time744.9 seconds × $0.08 / hour
$0.0166

Total, rounded to cents $1.81

Fable 5.1$5.33

Other input86 tokens × $10.00 / million
$0.0009
Save input100,669 tokens × $12.50 / million
$1.2584
Reuse input2,491,809 tokens × $0.25 / million
$0.6230
Model output68,554 tokens × $50.00 / million
$3.4277
Active session time1033.7 seconds × $0.08 / hour
$0.0230

Total, rounded to cents $5.33

Token prices saved on 3 September; active-time rate verified on 5 September. Charges are shown to four decimals; the complete calculation is rounded once to cents. Source: recorded token counts, prices, and active seconds.

Evidence limits

Model names identify configured sessions. Individual request records do not identify the serving model, so calculations apply each session’s saved model price to its recorded usage. Complete request inputs were not preserved. Exact reuse of particular instructions or tool results is unknown. The main page’s request-loop and caching diagrams explain the process; they do not reconstruct a recorded request.

Two runs in a fixed order cannot establish typical variation, causes for every action, or future savings. A grouped command can contain several operations, and a response can request several tools. Neither grouping nor preview counts establishes simultaneous execution, effectiveness, or complete review.

The test software did not explicitly set the effort control, which influences token use. Records show high effort for four settings and no value for Haiku. Three early Run 2 sessions had spending limits that differed from the plan. Effects beyond the recorded Sonnet email stop were not measured. The downloaded procedure documents these setup differences.

Future research will test common assumptions about applied AI and quantify the effects of recommended practices. We will compare the quality, usage cost, and completion time of the same work under different setups. Skills, templates, independent checks, MCP and API access, tool-call efficiency, model choice, conversation history, and caching are planned comparisons. Their effects on quality, cost, time, and consistency remain to be measured.

Original materials

The package’s offline checks need Python, not a model account. New experiments require compatible service and model access. The public records shorten long excerpts and replace private locations. Original outputs remain unchanged.

Both runs

Recorded task costs in US dollars. Review and stopping qualifications above apply to these values.

Forecast email
ModelRun 1Run 2
Haiku 4.5$0.09*$0.12*
Sonnet 5$0.26$0.36*
Opus 5$0.46$0.48
Fable 5$0.63$0.74
Fable 5.1$1.09*$1.29*

* See the numerical, length, or spending-limit qualification above.

One-pager
ModelRun 1Run 2
Haiku 4.5$0.19$0.24
Sonnet 5$1.81$1.56
Opus 5$4.94$2.86
Fable 5$2.59$3.32
Fable 5.1$5.33$5.06

* See the numerical, length, or spending-limit qualification above.

Slide deck
ModelRun 1Run 2
Haiku 4.5$0.33*$0.31
Sonnet 5$3.78$2.85
Opus 5$5.95$7.81
Fable 5$7.67$5.57
Fable 5.1$5.32$4.16

* See the numerical, length, or spending-limit qualification above.

Inspect all original outputs ↗

Official references

Mechanism and company references checked on 6 September 2026. Experiment calculations retain the saved study prices.

Further research

Uber’s software-factory research breaks spending into model requests, token use, and prices. Its examples compare tool loading, repeated status checks, and returning smaller results. Ramp’s Inspect reports describe giving agents a working environment for tests and visual checks, then directing them at specific failures.

These company reports inform future comparisons. Their reported savings and performance do not measure the results of this study.

Company-authored engineering reports, accessed 6 September 2026.

Source inventory: 30 original output files, with hashes preserved in the study package.