Tenex Research · Model selection
AI Model Evaluation: Comparing Quality, Cost, and Time
Five Claude models on the same sales assignments
At Tenex, we help companies put AI into everyday work, from enterprise rollouts to custom software. Clients adopting Claude Enterprise increasingly ask how to control usage costs without lowering the standard of the work. This study establishes a baseline for that analysis and builds a framework for robust future evaluation.
We gave five Claude models identical fictional sales records and instructions for a forecast email, a One-pager, and a slide deck. We compared the finished files, their cost and time, and the steps each model took. In future research, we will empirically evaluate commonly accepted beliefs around applied AI and quantify the impact of best practices across common knowledgework-flows.
How the loop works
An agentic loop is a repeated exchange: an application sends information to a model, and the model sends back text or a request to use a tool.
For example, Claude might ask a tool to read a sales file. The application runs the tool, then sends the result and the conversation so far back to Claude. Claude responds again, either requesting another action or returning an answer.
How requests repeat
Request input
Standing instructions (the system prompt) Your assignment Tool descriptions Earlier messages and results. Input token charges.
Model output
Claude returns an answer or asks a tool to act. A tool call can contain code or instructions. Output token charges.
Tool result
The application runs the tool. It returns data, an image, or an error message. No recorded tool fee.
The tool result can join the next request. When ready, Claude returns the answer or finished file.
What we measured
- Quality
- Selected figure checks, word limits, and visible formatting or wording defects. No overall quality score.
- Cost
- Recorded dollars: charges for model input and output, plus billed active session time.
- Time
- Minutes from the first task message to the recorded stop. Billed active time is measured separately.
- Requests and tools
- Model exchanges and the actions they requested. One exchange can ask for several tool calls.
- Tokens
- Units of model input and output, including text pieces and image input. Different categories have different prices.
Study setup
The data described 16 fictional sales opportunities at two dates, including deal values, forecast categories, owners, expected close dates, and activity dates. We also supplied sales targets, seller notes, and rules for calculating the forecast.
Each prompt asked for the forecast, an explanation of what changed, risks, and management actions. The models had to use the supplied files, show calculation inputs, cite evidence, and flag missing information. The email had to stay under 350 words, the One-pager fit one page, and the deck span six to eight pages.
The prompts, data, and review details are in Methods and sources.
Key findings
Choose by assignment
Sonnet’s One-pager and deck cost less than Opus’s. Fable 5.1 cost 8% more than Opus on the One-pager, but 11% less on the deck. Haiku was cheapest but gave the wrong forecast in its deck. Check the finished file before choosing on cost.
Test templates before cutting checks
Requests to build or check the One-pagers accounted for 56–92% of token charges. Test whether a fixed template reduces the work needed to produce the page. These are whole request bills; the share does not measure how much a template would save.
Token counts can mislead
Token charges cover the information sent to a model and the output it generates. Output costs more per token. In Fable 5.1’s One-pager, output accounted for about 3% of tokens but 65% of token charges. Compare spending on input and output before deciding where to reduce usage.
Quality
Correct figures were only one part of quality. Some files reported correct headline figures but gave misleading explanations; others used small text or crowded layouts. Review the calculations, the wording, and the finished page separately.
Sonnet, Opus, and both Fable versions produced more developed layouts than Haiku in the displayed files. Opus made the weighted calculation more explicit; Fable 5 separated movement, concentration, and actions clearly. Fable 5.1 added visual comparisons, with smaller text across the page. Haiku’s deck had an incorrect forecast and a clipped heading.
For the One-pager, Sonnet is a lower-cost candidate to test: it cost less than Opus and both Fable versions. Require correct explanations and a readable layout as well as correct figures. The numerical check alone did not establish that the file was ready to use.
We inspected the displayed files with model names visible; no overall score was assigned. A separate check of six headline figures per file found four incorrect values across two Haiku outputs. Three emails missed their word limits.
Cost
Sonnet’s One-pager cost $1.81, against $4.94 for Opus and $5.33 for Fable 5.1. On the deck, Fable 5.1 cost $5.32 against Opus’s $5.95. The same model could be the more expensive choice for one assignment and the cheaper choice for another.
Cost and completion time
Forecast email
One-pager
Slide deck
Prices saved on 3 September 2026. Each assignment has its own dollar scale. * Haiku’s and Fable 5.1’s emails missed the word limit. Haiku’s deck has three incorrect headline figures. Source: original task cost, time, and output-check records.
Time
A faster result can cost more. Fable 5’s One-pager took 5.9 minutes and cost $2.59; Sonnet’s took 12.4 minutes and cost $1.81. Compare models against the time available for the assignment, alongside the standard the file must meet.
Completion time by assignment
| Model | Run 1 | Run 2 |
|---|---|---|
| Haiku 4.5 | 1.8* | 2.3‡ |
| Sonnet 5 | 2.8 | 3.3† |
| Opus 5 | 2.4 | 2.2 |
| Fable 5 | 2.8 | 1.8 |
| Fable 5.1 | 4.9* | 5.2* |
| Model | Run 1 | Run 2 |
|---|---|---|
| Haiku 4.5 | 3.0 | 4.4 |
| Sonnet 5 | 12.4 | 12.2 |
| Opus 5 | 18.2 | 12.7 |
| Fable 5 | 5.9 | 7.5 |
| Fable 5.1 | 17.2 | 16.2 |
| Model | Run 1 | Run 2 |
|---|---|---|
| Haiku 4.5 | 6.7‡ | 6.2 |
| Sonnet 5 | 19.9 | 16.0 |
| Opus 5 | 22.4 | 24.9 |
| Fable 5 | 15.8 | 11.3 |
| Fable 5.1 | 15.9 | 13.0 |
* These emails missed the word limit. ‡ The marked Haiku email had one incorrect headline value; the marked deck had three. † Sonnet’s email stopped at its spending limit after saving the file.
Elapsed time includes model exchanges, tool work, and waits. Individual tool times were not measured separately. Source: original task timestamps and output checks.
Requests and tools
To produce the One-pagers, the models asked tools to read sales records, calculate figures, build the page, and check it. We grouped their requests by these activities and added up the token charges for each group.
Document requests cost the most
Create or edit: The model asks tools to write, edit, or turn a document into a PDF or preview.
The underline marks the selected activity. Charges cover the full model request, including input and output. Active session time is separate. Source: reviewed calls and request usage.
Sonnet’s 31 document-inspection requests cost $0.67, almost the same as Opus’s 10 requests at $0.66. A request count alone does not reveal the bill. One response can ask for several tools, and one tool call can contain several operations.
Sonnet opened its first One-pager preview at call 22, then made 62 more calls. These included measuring text, changing the page, rebuilding it, and opening cropped previews. Requests after that preview accounted for 75% of its token charges. Opus checked for overflowing text; Fable 5 made a shorter sequence of label and source-note changes.
Fable 5.1 made 41 One-pager calls against Fable 5’s 28. Its program expected nine deals with at least 60 days in their stage; the calculation found eight and the program stopped. It corrected the count, then repaired overflowing text and labels.
Test a fixed template for the page-building work, while keeping checks of the figures and final text. Measure the resulting charges: the document share identifies requests to investigate, not a saving we have demonstrated.
Tokens
Tokens measure model input and output. Text becomes pieces such as words, parts of words, and punctuation; images also count toward input usage. Output includes the answer, code sent to tools, and billed reasoning that may not appear in the returned text.
The application adds the model’s reply and tool results to the conversation before sending the next request. When earlier history is kept, requests grow. Prompt caching lets an unchanged beginning be reused at a lower input price, provided the saved copy is still available. Saving and reusing input have separate charges.
Cached request input
With earlier history retained
Sonnet 5 · One-pager · Run 1
Input and request costs
+ Active session time$0.0166
Recorded task cost$1.81
By request 6, input was 2.6 times its size on request 1. Cached material still counts as input, at a lower price.
Each request’s cost includes input and output. The task total adds active session time and rounds once to cents. Yellow includes input written to cache and other uncached input. Source: recorded request usage and saved prices.
Sonnet’s One-pager accumulated 5.07 million reused input tokens across 79 requests, costing $1.01. Reuse was cheaper per token, but each exchange added to the bill. Complete request inputs were not saved, so the exact text reused is unknown.
Token charges
Generating output and reusing input carry different prices. For Fable 5.1’s One-pager, output cost $50 per million tokens; reused input cost $0.25.
Most charges came from output
At those rates, one output token cost as much as 200 reused input tokens.
Output includes intermediate responses, code sent to tools, and billed reasoning. Percentages use token charges only; active session time adds about $0.02. Source: Fable 5.1’s recorded One-pager usage and prices.
Output accounted for most spending in this example. Test whether a document template reduces generated code while keeping the finished work accurate and readable.
Future research
Future research will test common assumptions about applied AI and quantify the effects of recommended practices. We will compare the quality, usage cost, and completion time of the same work under different setups.
Skills provide reusable instructions and files for a task. We will test if skills can reduce usage costs and make outputs more consistent. Reusable templates and independent checks will be evaluated in the same way, including whether additional checks improve the work enough to justify their cost.
Other comparisons will examine how models access information: through MCP tools or direct API calls, with more targeted searches and smaller results. MCP connects AI applications to tools; an API lets software request data or actions. MCP tools can themselves call APIs. We will also test model choice, how much reasoning a model performs, conversation history, caching, and when to stop checks or retries.
