# Editorial evidence

Review date: 6 September 2026. Report: How AI Models Approach Work.

## Review scope

The first-attempt emails were read in full. All five first-attempt briefs and every page of the five first-attempt decks were rendered and visually inspected. Model names were visible. These are specific design, formatting, and wording observations, not a blind quality score or a comprehensive validation. Repeated outputs remain available; the new visual comments cover the first attempt.

The separate [numerical review](output-headline-audit.md) checked six selected figures per output. It found four incorrect values in two Haiku outputs. Additional wording and numerical observations below are outside that six-figure count. Three emails missed their word limits. Sonnet’s repeated email stopped at its spending limit after saving its unchanged file; displayed totals end at that stop.

## Haiku’s category mismatch

Sources: first Haiku deck, calls 11–12; final deck page 2; supplied current-pipeline.csv, prior-pipeline.csv, and forecast-rules.md. These files are in the [study package](replication.zip). The [public session record](session-details/run_one-a-haiku.json) contains shortened instructions and the returned forecast values.

The complete original command was also inspected because its public excerpt ends before the lookup. The [relevant original code](haiku-category-excerpt.py) is provided separately. It assigns 60% to `Best Case` and uses `weights.get(row['ForecastCategoryName'], 0)`. The source category is `BestCase`. The lookup therefore returns zero for that category.

Current BestCase amounts total $950,000. Their contribution should be $570,000. Correct calculation gives $1,330,000; the mismatched calculation gives $760,000. The returned record and final deck contain $760,000. This reproduces the error from the supplied data. It does not establish a cause for unrelated model behavior.

The original commands read the rules and source files. Later calls created preview images but contain no recorded image-opening calls. Creating an image does not show that the model inspected it. The first brief did record two preview openings. Haiku’s repeated deck matched all six selected headline figures.

Original source: `runs/managed-agents-pilot-2026-09-03-v2/a-haiku-r1/tool-calls-task.jsonl`.
Original source SHA-256: `ca3bb3d234960452ec4298428c1add4506cb55b410562860ee4cdd1c6432ac8d`.

## Sonnet’s wording edit

Sources: [first Sonnet brief record](session-details/run_one-c-sonnet.json), calls 44, 48, and 51; final brief, concentration section.

Call 44 measured a line containing both unweighted and weighted deal amounts: 427.986 points, with 325 available. Call 48 measured shorter candidates. Call 51 replaced the line with:

> Dovetail Travel $420K + Aster Health $340K = $624K weighted (47%)

The prior wording named Aster’s weighted $204K contribution within the line. The replacement leaves that value only on the following line. Aster’s 60% weight is required to turn $340K into $204K. The total is correct; the shortened equation is misleading in isolation. This is a visible wording defect, separate from the selected numerical audit.

Sonnet’s first brief used 84 calls. The first preview-opening call was 22, leaving 62 subsequent calls. They include measurements, edits, rebuilding, and preview openings. Sonnet’s first deck was built incrementally, so activity after its first preview also includes creating later slides.

## Fable’s additional work

Sources: [Fable 5 brief](session-details/run_one-c-fable.json), [Fable 5.1 brief](session-details/run_one-c-fable-5.1.json), and corresponding deck and email records.

First-attempt tool calls:

| Assignment | Fable 5 | Fable 5.1 |
| --- | ---: | ---: |
| Forecast email | 10 | 15 |
| One-page brief | 28 | 41 |
| Slide deck | 71 | 41 |

Fable 5.1’s brief call 8 failed a check expecting nine deals at least 60 days into their stage. Call 9 returned eight and listed the deals. Call 10 changed the expected count from nine to eight, then encountered a separate movement-section overflow. Later calls adjusted heights, text, and labels. Those extra calls include repair and formatting.

Fable 5’s first deck contains 26 direct text-edit calls to its deck program. Fable 5.1 often combined editing and rebuilding within commands. Call totals count different amounts of work and do not establish simultaneous execution or effectiveness.

Fable 5.1’s first brief generated 68,554 output tokens against Fable 5’s 26,110. Both recorded output prices were $50 per million. Reuse prices were $0.25 and $1.00 per million respectively. The costs were $5.33 and $2.59. Fable 5.1’s first deck instead cost $5.32 against Fable 5’s $7.67.

## Charges after the preview

The [derived values](editorial-findings.json) join the [tool records](exploration-data.json) to [request charges](cost-exploration-data.json). The boundary is the first call whose reviewed action begins `Open the image`. Later charges sum requests with a strictly higher request number than the request containing that call. The boundary request stays in the earlier segment.

| First brief | Token charges | Later share |
| --- | ---: | ---: |
| Haiku 4.5 | $0.18 | 37% |
| Sonnet 5 | $1.80 | 75% |
| Opus 5 | $4.91 | 58% |
| Fable 5 | $2.58 | 35% |
| Fable 5.1 | $5.31 | 36% |

These are token charges rounded for display, excluding active session time. The later share covers all subsequent activity, not revision alone. Tools have no separate dollar charge or exclusive measured execution time.

Sonnet’s first brief reused 5,067,826 input tokens at $0.20 per million: $1.0135652, about 56% of token charges. Fable 5.1’s brief generated 68,554 output tokens at $50 per million: $3.4277, about 65% of token charges and 3% of tokens. Full unrounded category calculations plus active time reproduce the recorded costs.

Fable 5.1’s first brief cost 7.9% more than Opus’s ($5.33 versus $4.94); its first deck cost 10.6% less ($5.32 versus $5.95). These task-specific comparisons are not expected future savings.

## Original output observations

### Haiku 4.5: email

Source: [original file](/outputs/managed-agents/run-one/email--haiku-4.5.md).

Section headings organize the email, but long paragraphs, deal IDs, and repeated source references make it dense. The body contains 530 words against a limit of fewer than 350.

The six checked headline figures match the source calculations. The movement paragraph assigns Lantern’s amount increase to Junction. The repeated email corrected the length but reported the top-two share as 57% instead of 47%.

### Haiku 4.5: brief

Source: [original file](/outputs/managed-agents/run-one/one-page-pdf--haiku-4.5.pdf), pages 1.

Large blank areas surround small action text. The first panel has substantial unused height, while the movement note crowds the next heading. Square replacement characters appear in the actions.

The six checked headline figures agree with the source. The movement section says four categories upgraded; three categories changed and a fourth deal increased in amount.

### Haiku 4.5: deck

Source: [original file](/outputs/managed-agents/run-one/deck--haiku-4.5.pdf), pages 2, 3, 6.

Several slides use only their upper half. Page 3 places a small table in a mostly empty page. Page 6 clips the long title at the right edge, and text from the two columns overlaps.

Page 2 reports a $760,000 forecast instead of $1,330,000, a 61% concentration instead of 47%, and a $276K increase instead of $384K. These errors arose in the saved calculation before slide creation.

### Sonnet 5: email

Source: [original file](/outputs/managed-agents/run-one/email--sonnet-5.md).

Bold labels separate the sections within a 290-word body. The movement paragraph compresses several deal IDs and calculations into one block.

The six checked headline figures agree. The email says the increase was driven entirely by three deals, then also lists Lantern’s separate $6,250 contribution.

### Sonnet 5: brief

Source: [original file](/outputs/managed-agents/run-one/one-page-pdf--sonnet-5.pdf), pages 1.

Large headline figures establish a clear entry point. Two columns organize the detail, but the movement and concentration sections are dense; the source footer is smaller and lighter.

The six headline figures agree. In the concentration section, $420K plus $340K is shown as $624K weighted without applying Aster’s weight in that line. The following line correctly gives Aster’s $204K weighted value.

### Sonnet 5: deck

Source: [original file](/outputs/managed-agents/run-one/deck--sonnet-5.pdf), pages 3, 4, 7.

The eight slides use a consistent teal-and-orange palette and finding-led headings. Page 3 aligns owner bars with a data table. The page-4 movement chart and page-7 action table contain much smaller supporting text.

The six checked headline figures agree. Page 4 says the gain is entirely from three category changes, while its own chart includes a separate $6,250 amount increase.

### Opus 5: email

Source: [original file](/outputs/managed-agents/run-one/email--opus-5.md).

Bold section openings organize a 332-word body. Calculations and source references remain embedded in the paragraphs, with the three actions in one continuous block.

The six checked headline figures agree. The email distinguishes the $378,000 category contribution from the $384,250 total increase and identifies backward stage-entry dates.

### Opus 5: brief

Source: [original file](/outputs/managed-agents/run-one/one-page-pdf--opus-5.pdf), pages 1.

Four metric panels lead into aligned tables. The concentration table shows each deal’s weighted value directly. The page fits more comparisons and source qualifications than Sonnet’s, with smaller text in the lower sections.

The six checked headline figures agree. The concentration calculation is explicit: $420,000 plus $204,000 equals $624,000. No overall quality score was assigned.

### Opus 5: deck

Source: [original file](/outputs/managed-agents/run-one/deck--opus-5.pdf), pages 2, 4, 7.

A navy rule, orange annotations, and repeated heading positions unify the eight slides. Page 2 separates the forecast calculation from its explanation. Pages 4 and 7 rely on dense tables and small explanatory notes.

The six checked headline figures agree. Page 4 says five deals were re-rated, but its table shows three category changes, one amount change, and one date-only change.

### Fable 5: email

Source: [original file](/outputs/managed-agents/run-one/email--fable-5.md).

The 305-word body separates actions into a numbered list. Section openings make the forecast and data issues easy to locate, though the movement paragraph contains many compact figures.

The six checked headline figures agree. The email separates category changes from the amount change and names missing target-period information.

### Fable 5: brief

Source: [original file](/outputs/managed-agents/run-one/one-page-pdf--fable-5.pdf), pages 1.

A large metric band and three aligned columns separate movement, concentration, and actions. Small bars show the two largest deals’ contributions. Supporting text remains dense, particularly in the movement column.

The six checked headline figures agree. The movement section explicitly separates the $378,000 category contribution and $6,250 amount contribution.

### Fable 5: deck

Source: [original file](/outputs/managed-agents/run-one/deck--fable-5.pdf), pages 3, 4, 7.

The eight-slide deck consistently pairs large finding headings with charts and tables. Page 4 gives the movement chart room beside deal-level changes. Blue source notes and orange decision notes recur in stable positions.

The six checked headline figures agree. Page 4 correctly attributes about 98% of growth to category upgrades, rather than all growth. Its $7.67 cost was the highest first-attempt deck cost.

### Fable 5.1: email

Source: [original file](/outputs/managed-agents/run-one/email--fable-5.1.md).

Numbered actions and a separate data-problems list break up the text. The final body has 351 words and misses the fewer-than-350 requirement despite repeated shortening.

The six checked headline figures agree. It identifies activity dates after the prior snapshot. It also says Morgan has five opportunities; the current source file lists four. That additional observation is separate from the six-figure audit.

### Fable 5.1: brief

Source: [original file](/outputs/managed-agents/run-one/one-page-pdf--fable-5.1.pdf), pages 1.

The top panel compares open and weighted values with stacked bars. Owner coverage and inactive deals have separate tables. This gives more visual comparisons than Fable 5’s brief, but packs small text into nearly every section.

The six checked headline figures agree. The page distinguishes category and amount contributions. During creation, the model corrected its own expected count of aged deals from nine to eight.

### Fable 5.1: deck

Source: [original file](/outputs/managed-agents/run-one/deck--fable-5.1.pdf), pages 4, 5, 6.

Consistent navy headings and orange annotations connect all eight slides. Page 5 combines concentration bars, a deal table, and a horizontal age chart. Page 6 devotes most of the slide to a detailed table; full-size viewing is needed for its notes.

The six checked headline figures agree. Page 4 attributes all $384,250 of growth to three category changes, although the same page separately shows $378,000 from categories and $6,250 from amount.

## Subsequent experiments

Skills, templates, independent calculation checks, and blind output review are proposed next comparisons. Their effect on output quality, cost, and time is untested. They are motivated by the specific defects above; no savings or quality improvement is claimed in advance.

## Mechanism references

General request-loop and caching explanations follow official Claude documentation, checked 6 September 2026:

- [Tool use](https://platform.claude.com/docs/en/agents-and-tools/tool-use/how-tool-use-works)
- [Prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching)
- [Token counting](https://platform.claude.com/docs/en/build-with-claude/token-counting)
- [Reasoning](https://platform.claude.com/docs/en/build-with-claude/thinking)

Complete request inputs were not preserved. Exact reuse of particular text is unknown. Session labels identify configured models; individual requests do not identify the serving model. Two fixed-order attempts do not establish permanent model traits.


## Request activity review

The September 6 refinement groups whole request charges from the five Run 1 briefs by their requested tool operations. The review covers all 236 calls in those briefs. The activity records include source indexes, full-input hashes, definitions and every request’s original charge.

Document requests account for 56–92% of token charges across the five briefs. Sonnet’s 31 document-inspection requests cost $0.67101000; its 30 file-building requests cost $0.69143090. Reused input accounted for 71.8% of the inspection group. Fable 5.1’s file-building requests cost $3.20823000; model output contributed 78.7%.

These groups include the entire request input and output and exclude active session time. Mixed actions remain together. They do not price individual tools, measure formatting time or establish removable costs. The next study will test skills for lower usage costs and more consistent outputs; those benefits have not yet been measured.

Sources: [activity charges](./activity-costs.json), [reviewed call labels](./brief-activity-review.json), and [original request charges](./cost-exploration-data.json).
