Tenex Research · Model selection

AI Model Evaluation: Comparing Quality, Cost, and Time

Five Claude models on the same sales assignments

At Tenex, we help companies put AI into everyday work, from enterprise rollouts to custom software. Clients adopting Claude Enterprise increasingly ask how to control usage costs without lowering the standard of the work. This study establishes a baseline for that analysis and builds a framework for robust future evaluation.

We gave five Claude models identical fictional sales records and instructions for a forecast email, a One-pager, and a slide deck. We compared the finished files, their cost and time, and the steps each model took. In future research, we will empirically evaluate commonly accepted beliefs around applied AI and quantify the impact of best practices across common knowledgework-flows.

How the loop works

An agentic loop is a repeated exchange: an application sends information to a model, and the model sends back text or a request to use a tool.

For example, Claude might ask a tool to read a sales file. The application runs the tool, then sends the result and the conversation so far back to Claude. Claude responds again, either requesting another action or returning an answer.

How requests repeat

Request input

Standing instructions (the system prompt) Your assignment Tool descriptions Earlier messages and results. Input token charges.

Model output

Claude returns an answer or asks a tool to act. A tool call can contain code or instructions. Output token charges.

Tool result

The application runs the tool. It returns data, an image, or an error message. No recorded tool fee.

The tool result can join the next request. When ready, Claude returns the answer or finished file.

What we measured

Quality
Selected figure checks, word limits, and visible formatting or wording defects. No overall quality score.
Cost
Recorded dollars: charges for model input and output, plus billed active session time.
Time
Minutes from the first task message to the recorded stop. Billed active time is measured separately.
Requests and tools
Model exchanges and the actions they requested. One exchange can ask for several tool calls.
Tokens
Units of model input and output, including text pieces and image input. Different categories have different prices.

Study setup

The data described 16 fictional sales opportunities at two dates, including deal values, forecast categories, owners, expected close dates, and activity dates. We also supplied sales targets, seller notes, and rules for calculating the forecast.

Each prompt asked for the forecast, an explanation of what changed, risks, and management actions. The models had to use the supplied files, show calculation inputs, cite evidence, and flag missing information. The email had to stay under 350 words, the One-pager fit one page, and the deck span six to eight pages.

The prompts, data, and review details are in Methods and sources.

Key findings

Choose by assignment

Sonnet’s One-pager and deck cost less than Opus’s. Fable 5.1 cost 8% more than Opus on the One-pager, but 11% less on the deck. Haiku was cheapest but gave the wrong forecast in its deck. Check the finished file before choosing on cost.

Test templates before cutting checks

Requests to build or check the One-pagers accounted for 56–92% of token charges. Test whether a fixed template reduces the work needed to produce the page. These are whole request bills; the share does not measure how much a template would save.

Token counts can mislead

Token charges cover the information sent to a model and the output it generates. Output costs more per token. In Fable 5.1’s One-pager, output accounted for about 3% of tokens but 65% of token charges. Compare spending on input and output before deciding where to reduce usage.

Quality

Correct figures were only one part of quality. Some files reported correct headline figures but gave misleading explanations; others used small text or crowded layouts. Review the calculations, the wording, and the finished page separately.

Sonnet, Opus, and both Fable versions produced more developed layouts than Haiku in the displayed files. Opus made the weighted calculation more explicit; Fable 5 separated movement, concentration, and actions clearly. Fable 5.1 added visual comparisons, with smaller text across the page. Haiku’s deck had an incorrect forecast and a clipped heading.

Recorded cost
$1.81
Completion time
12.4 min
Sonnet 5 · One-pagerOne page
Sonnet 5, One-pager, Run 1, original page 1 of 1

Large forecast figures appear at the top. Two columns organize the detail, but the movement and concentration sections are dense; the source footer is smaller and lighter.

The six headline figures agree. In the concentration section, $420K plus $340K is shown as $624K weighted without applying Aster’s weight in that line. The following line correctly gives Aster’s $204K weighted value.

For the One-pager, Sonnet is a lower-cost candidate to test: it cost less than Opus and both Fable versions. Require correct explanations and a readable layout as well as correct figures. The numerical check alone did not establish that the file was ready to use.

We inspected the displayed files with model names visible; no overall score was assigned. A separate check of six headline figures per file found four incorrect values across two Haiku outputs. Three emails missed their word limits.

Cost

Sonnet’s One-pager cost $1.81, against $4.94 for Opus and $5.33 for Fable 5.1. On the deck, Fable 5.1 cost $5.32 against Opus’s $5.95. The same model could be the more expensive choice for one assignment and the cheaper choice for another.

Run 1 · US dollars

Cost and completion time

Forecast email

Model / costMinutes
Haiku 4.5*$0.091.8
Sonnet 5$0.262.8
Opus 5$0.462.4
Fable 5$0.632.8
Fable 5.1*$1.094.9

One-pager

Model / costMinutes
Haiku 4.5$0.193.0
Sonnet 5$1.8112.4
Opus 5$4.9418.2
Fable 5$2.595.9
Fable 5.1$5.3317.2

Slide deck

Model / costMinutes
Haiku 4.5*$0.336.7
Sonnet 5$3.7819.9
Opus 5$5.9522.4
Fable 5$7.6715.8
Fable 5.1$5.3215.9

Prices saved on 3 September 2026. Each assignment has its own dollar scale. * Haiku’s and Fable 5.1’s emails missed the word limit. Haiku’s deck has three incorrect headline figures. Source: original task cost, time, and output-check records.

Time

A faster result can cost more. Fable 5’s One-pager took 5.9 minutes and cost $2.59; Sonnet’s took 12.4 minutes and cost $1.81. Compare models against the time available for the assignment, alongside the standard the file must meet.

Elapsed minutes · Both runs

Completion time by assignment

Forecast email
ModelRun 1Run 2
Haiku 4.51.8*2.3
Sonnet 52.83.3
Opus 52.42.2
Fable 52.81.8
Fable 5.14.9*5.2*
One-pager
ModelRun 1Run 2
Haiku 4.53.04.4
Sonnet 512.412.2
Opus 518.212.7
Fable 55.97.5
Fable 5.117.216.2
Slide deck
ModelRun 1Run 2
Haiku 4.56.76.2
Sonnet 519.916.0
Opus 522.424.9
Fable 515.811.3
Fable 5.115.913.0

* These emails missed the word limit. ‡ The marked Haiku email had one incorrect headline value; the marked deck had three. † Sonnet’s email stopped at its spending limit after saving the file.

Elapsed time includes model exchanges, tool work, and waits. Individual tool times were not measured separately. Source: original task timestamps and output checks.

Requests and tools

To produce the One-pagers, the models asked tools to read sales records, calculate figures, build the page, and check it. We grouped their requests by these activities and added up the token charges for each group.

One-pager · Run 1

Document requests cost the most

Create or edit: The model asks tools to write, edit, or turn a document into a PDF or preview.

ModelTotal token chargesCreate or edit
Haiku 4.5
Read or calculate: $0.0399 · 11 requestsCreate or edit: $0.0749 · 5 requestsCheck files: $0.0477 · 9 requestsCreate + check: $0.0000 · 0 requestsOther work: $0.0103 · 1 requestsNo tool call: $0.0085 · 1 requests$0.18
$0.075 requests
Sonnet 5
Read or calculate: $0.1780 · 9 requestsCreate or edit: $0.6914 · 30 requestsCheck files: $0.6710 · 31 requestsCreate + check: $0.0577 · 3 requestsOther work: $0.1706 · 5 requestsNo tool call: $0.0290 · 1 requests$1.80
$0.6930 requests
Opus 5
Read or calculate: $0.4405 · 8 requestsCreate or edit: $3.0492 · 29 requestsCheck files: $0.6612 · 10 requestsCreate + check: $0.0495 · 1 requestsOther work: $0.6137 · 6 requestsNo tool call: $0.1007 · 1 requests$4.91
$3.0529 requests
Fable 5
Read or calculate: $0.4086 · 7 requestsCreate or edit: $1.2112 · 11 requestsCheck files: $0.2264 · 4 requestsCreate + check: $0.0000 · 0 requestsOther work: $0.6257 · 2 requestsNo tool call: $0.1098 · 1 requests$2.58
$1.2111 requests
Fable 5.1
Read or calculate: $0.2391 · 6 requestsCreate or edit: $3.2082 · 15 requestsCheck files: $1.0642 · 14 requestsCreate + check: $0.5998 · 5 requestsOther work: $0.0962 · 1 requestsNo tool call: $0.1024 · 1 requests$5.31
$3.2115 requests

The underline marks the selected activity. Charges cover the full model request, including input and output. Active session time is separate. Source: reviewed calls and request usage.

Sonnet’s 31 document-inspection requests cost $0.67, almost the same as Opus’s 10 requests at $0.66. A request count alone does not reveal the bill. One response can ask for several tools, and one tool call can contain several operations.

Sonnet opened its first One-pager preview at call 22, then made 62 more calls. These included measuring text, changing the page, rebuilding it, and opening cropped previews. Requests after that preview accounted for 75% of its token charges. Opus checked for overflowing text; Fable 5 made a shorter sequence of label and source-note changes.

Fable 5.1 made 41 One-pager calls against Fable 5’s 28. Its program expected nine deals with at least 60 days in their stage; the calculation found eight and the program stopped. It corrected the count, then repaired overflowing text and labels.

Test a fixed template for the page-building work, while keeping checks of the figures and final text. Measure the resulting charges: the document share identifies requests to investigate, not a saving we have demonstrated.

Tokens

Tokens measure model input and output. Text becomes pieces such as words, parts of words, and punctuation; images also count toward input usage. Output includes the answer, code sent to tools, and billed reasoning that may not appear in the returned text.

The application adds the model’s reply and tool results to the conversation before sending the next request. When earlier history is kept, requests grow. Prompt caching lets an unchanged beginning be reused at a lower input price, provided the saved copy is still available. Saving and reusing input have separate charges.

Cached request input

With earlier history retained

Earlier inputModel replyTool resultsNext request input

Sonnet 5 · One-pager · Run 1

Input and request costs

Reused from cacheNot read from cache
RequestInput tokensInput costOutput costTotal cost
Request 1
6,459
$0.0161$0.0025$0.0186
Request 2
7,812
$0.0047$0.0016$0.0063
Request 3
8,651
$0.0037$0.0020$0.0056
Request 4
9,004
$0.0026$0.0041$0.0067
Request 5
9,986
$0.0043$0.0030$0.0072
Request 6
16,693
$0.0188$0.0168$0.0356
Requests 779$1.2211$0.4964$1.7175
All 79 requests$1.2712$0.5265$1.7977

+ Active session time$0.0166

Recorded task cost$1.81

By request 6, input was 2.6 times its size on request 1. Cached material still counts as input, at a lower price.

Each request’s cost includes input and output. The task total adds active session time and rounds once to cents. Yellow includes input written to cache and other uncached input. Source: recorded request usage and saved prices.

Sonnet’s One-pager accumulated 5.07 million reused input tokens across 79 requests, costing $1.01. Reuse was cheaper per token, but each exchange added to the bill. Complete request inputs were not saved, so the exact text reused is unknown.

Token charges

Generating output and reusing input carry different prices. For Fable 5.1’s One-pager, output cost $50 per million tokens; reused input cost $0.25.

Fable 5.1 · One-pager · Run 1

Most charges came from output

Model outputAll input categories
Share of tokens3%

68,554 of 2,661,118

Share of token charges65%

$3.43 of $5.31

At those rates, one output token cost as much as 200 reused input tokens.

Output includes intermediate responses, code sent to tools, and billed reasoning. Percentages use token charges only; active session time adds about $0.02. Source: Fable 5.1’s recorded One-pager usage and prices.

Output accounted for most spending in this example. Test whether a document template reduces generated code while keeping the finished work accurate and readable.

Future research

Future research will test common assumptions about applied AI and quantify the effects of recommended practices. We will compare the quality, usage cost, and completion time of the same work under different setups.

Skills provide reusable instructions and files for a task. We will test if skills can reduce usage costs and make outputs more consistent. Reusable templates and independent checks will be evaluated in the same way, including whether additional checks improve the work enough to justify their cost.

Other comparisons will examine how models access information: through MCP tools or direct API calls, with more targeted searches and smaller results. MCP connects AI applications to tools; an API lets software request data or actions. MCP tools can themselves call APIs. We will also test model choice, how much reasoning a model performs, conversation history, caching, and when to stop checks or retries.