How to Read the Benchmark Charts in a Model Launch: Effort Curves, Log Cost Axes, and Safeguard Interventions
Model-launch accuracy-vs-cost charts carry hidden methodology choices. Worked on the 2026-09-01 Claude Fable 5.1 announcement: effort curves, log axes, error bars, scoring conventions, missing cells, safeguard zero-scores.
How to
Identify models and tiers
Map each curve and point to a model and an effort tier; lock the models you compare onto the same tier before reading numbers.
Convert the log cost axis
On 'log scale', convert visual distance into multiples: read the axis ends (e.g., $0.20-4 or $10-50) and estimate true multiples between points.
Check error bars and trials
Find the standard error and trial count in the footnotes (e.g., ±3.5-4.5 pts, 3 trials/task); ranking gaps inside the error range carry no conclusion.
Confirm the scoring convention
Distinguish partial vs strict and no-tools vs with-tools; cite the convention together with the score.
Handle missing cells and zero scores
Dashes mean unreported, not losing; check whether safeguards intervened and score zero, and read raw capability separately from the safeguarded product.
Separate in-house setups from public leaderboards
Publisher numbers and leaderboard numbers can legitimately differ; confirm harness and task-set version before choosing which to cite.
Read the methodology notes
Finish with the footnotes (task-set date, harness, reproduction conditions) and decide which external data this table can be compared against directly.
Model launches now ship with a standard chart type: accuracy-vs-cost scatter curves plus a multi-model comparison table. These charts are evidence -- but they embed a chain of methodology choices: effort tiers, scoring conventions, task-set versions, safeguard rules, error margins. The same model can move by tens of points across conventions. This article uses the 2026-09-01 Claude Fable 5.1 / Mythos 5.1 launch announcement as a worked example and unpacks eight reading points. All numbers are relayed from that announcement.
1. First: Effort Tiers Turn One Model into a Curve
Modern reasoning models have effort tiers -- in this announcement, low / med / high / xhigh / max, with official defaults: High in Claude Code, Medium in Claude Cowork and on Claude.ai. That means "a model" on an accuracy-vs-cost chart is a curve of tier points, not a single point.
Two immediate consequences:
- Cross-model comparisons must lock the same tier. The announcement's own claims are tier-scoped -- "at Low or Medium effort, Fable 5.1 achieves results similar to or better than Fable 5's at a much lower cost" says nothing about xhigh/max.
- Cross-vendor comparisons must also align tier semantics. GPT-5.6's reasoning.effort and this announcement's effort are both multi-tier designs, but each tier maps to different cost-quality trade-offs; read both vendors' tier documentation before comparing.
2. Log Cost Axes: Visual Distance Is Not Money Distance
Every cost x-axis in the announcement is labeled "USD, log scale" -- and the same page carries two ranges: $10-50 (Terminal-Bench-Science) and $0.20-4 (Humanity's Last Exam). On a log axis, equal spacing means equal multiples, not equal dollars; the visual midpoint of a range is its geometric mean.
Practice: before reading any point, read the axis's two ends and convert positions into multiples.
Three steps for reading the axis (HLE chart example):
1. Axis range: $0.20 -- $4 (log scale)
2. Visual midpoint ≈ geometric mean of $0.20 and $4 ≈ $0.89 -- not $2.1
3. Half the visual distance ≈ a multiple difference (read the ticks), not half the money
3. Error Bars and Trial Counts Are Hard Constraints
The announcement's footnote states Terminal-Bench-Science 0.1 carries a standard error of ±3.5-4.5 points per model, and the corresponding public leaderboard runs 3 trials/task (Claude Code harness). A ranking gap under 2 points therefore supports no conclusion.
The same page provides an official two-setups comparison: the public leaderboard reports Opus 5 at 30.0% and Fable 5 at 21.4%; the publisher's own setup reproduces 29.0% and 24.7% -- within noise, per the footnote. Note Fable 5's two numbers differ by 3.3 points: different harnesses and trial counts genuinely produce different numbers. Before citing any score, ask three questions: which task-set version, which harness, how many trials.
4. Scoring Conventions: partial vs strict Can Differ by 36 Points
OSWorld 2.0 (computer use) appears in the announcement's table under both conventions:
| Model | partial | strict |
|---|---|---|
| Fable 5.1 | 77.9% | 41.7% |
| Fable 5 | 72.9% | 36.1% |
| Opus 5 | 75.4% | 39.6% |
| GPT-5.6 Sol | — | — |
Same model, same task set, and the partial convention scores 36.2 points higher (calculated from the table). Partial typically credits partial completion; strict requires full success. Citing a computer-use score without its convention is an invalid citation. The paired convention also applies to HLE's no-tools / with-tools rows (Fable 5.1: 60.9% / 65.0%).
5. A Missing Cell (—) Is Unreported, Not Losing
The table has no GPT-5.6 Sol numbers for OSWorld 2.0 or Humanity's Last Exam -- the cells are dashes. A missing cell only means the publisher did not run or report that benchmark; reading it as "scored zero" or "avoided it" goes beyond the evidence. The correct move: cite an independent source for that benchmark, or state explicitly that the cell has no public data.
6. Tasks Where Safeguards Intervene Score Zero
The most commonly missed rule: the announcement states Fable 5.1 was evaluated with production safeguards enabled, and tasks where safeguards intervened scored zero -- on OSWorld 2.0 for both Fable tiers and on AutomationBench for Fable 5 -- which Anthropic acknowledges likely reduced those scores.
The same logic explains the two-tier spread: on Terminal-Bench 4.0, Mythos 5.1 (the stronger-safeguard tier) scores 60.9% vs Fable 5.1's 55.8%, officially attributed to earlier, less precise cyber safeguards, with the gap expected to shrink as safeguards improve. Keep "raw model capability" and "the safeguarded product" in separate reading lanes -- cite the one you would actually deploy.
7. Methodology Notes Decide Comparability
The announcement's OSWorld 2.0 footnote: scores are on the benchmark authors' August 2026 task release, with Fable 5 and Opus 5 re-run under the same conditions, and because the task files differ from earlier releases, these numbers are not directly comparable to previous ones. In other words, "77.9% on OSWorld 2.0" holds only within that task-set version; do not mix it with older OSWorld scores circulating online into one timeline table.
Convention checklist (run before citing any score):
[ ] Task-set / dataset version and date
[ ] Scoring convention (partial/strict, no-tools/with-tools)
[ ] Effort tier and default
[ ] Safeguards on? How are interventions scored?
[ ] Error margins and trial counts
[ ] Harness (evaluation framework) name
8. In Practice: Compressing the Table into Three Citable Sentences
Applying the eight points to this announcement's table yields three citable sentences:
# Two common quantities computed directly from the table:
# the partial-strict spread, and cost multiples at a locked tier
fable51_partial, fable51_strict = 77.9, 41.7
print(f"OSWorld 2.0 partial-strict spread: {fable51_partial - fable51_strict:.1f} pts")
# Cost multiple on a log axis: HLE chart range $0.20-$4, point A $0.5, point B $2.0
print(f"Same-tier cost multiple: {2.0 / 0.5:.1f}x (read the ticks, not the visual distance)")
Sentence one: under the official setup, Fable 5.1 leads on the five benchmarks with Sol scores, OSWorld and HLE report no Sol scores, and its results are with production safeguards enabled. Sentence two: every benchmark carries explicit error and convention limits (±3.5-4.5 pts, partial/strict, the 2026-08 task set); align conventions before mixing tables. Sentence three: effort defaults differ by product (Claude Code = High, Claude Cowork / Claude.ai = Medium); lock the tier for cross-vendor pricing. None of the three goes beyond the announcement plus arithmetic -- that is what "reading as evidence" means.
9. Common Mistakes and Troubleshooting
- Treating a low-tier point as "the model got stronger": cheaper tiers trade score for cost; cross-tier comparison is meaningless.
- Eyeballing multiples on a log axis: the visual midpoint approximates the geometric mean; eyeballing overestimates.
- Citing a partial score against someone else's strict score: a 36.2-point convention spread in this example.
- Reading a missing cell as a loss: a dash is "unreported".
- Ignoring safeguard-intervention zeros: the safeguarded product's score and the raw model's capability are two different quantities.
- Mixing task-set versions into one timeline: the methodology notes explicitly forbid it.
10. Next Steps
- The full fact briefing for this announcement: OpenAI Ecosystem September 2 Briefing: Anthropic Launches Claude Fable 5.1 / Mythos 5.1, Benchmarks Include GPT-5.6 Sol.
- The HLE benchmark itself (2,500 questions, published in Nature, dynamic HLE-Rolling): the official Humanity's Last Exam site.
- Applying this verification discipline to content production: Source Verification and Review-Until-Clean: A Fact-Checking Workflow for AI Content.
- The broader quality-control picture for deep research: DeepResearch Output Quality Control: Model Selection, Prompt Tuning, and Citation Verification.
Key points
- Effort tiers (low / med / high / xhigh / max) turn 'a model' into a cost-quality curve: cross-vendor price comparisons must lock the same tier or they do not hold
- Cost axes are almost always log scale: the same announcement page carries $0.20-4 and $10-50 x-axes, so a visual midpoint can hide a 5-10x gap
- Error bars are a hard constraint: Terminal-Bench-Science 0.1 carries ±3.5-4.5 pts standard error per model and the public leaderboard runs 3 trials/task
- Partial and strict are two different scores: OSWorld 2.0 shows Fable 5.1 at 77.9% partial vs 41.7% strict -- a 36.2-point spread; always cite the convention
- A missing cell (dash) means unreported, not losing: the announcement table has no GPT-5.6 Sol scores for OSWorld 2.0 or HLE
- Tasks where safety safeguards intervene score zero: Fable 5.1 ran with production safeguards, which Anthropic says likely depressed its scores; the Mythos-vs-Fable gap reflects safeguard granularity
- An in-house setup and the public leaderboard can differ: the leaderboard reports Opus 5 at 30.0% / Fable 5 at 21.4%, Anthropic's setup reproduces 29.0% / 24.7%, all within noise -- ask which harness and how many trials
- Methodology notes decide comparability: OSWorld 2.0 used the benchmark authors' August 2026 task release and is officially not directly comparable to earlier versions
Frequently asked questions
Official references
Related articles
Deep Research multi-model: GPT-5.6 vs Claude 4.5 vs Gemini 2.5 + third-party research tools
Deep Research across GPT-5.6 / Claude 4.5 / Gemini 2.5 - 6 real scenario tests + vs Elicit / Consensus / Perplexity. Which DR for which scenario, when to use third-party tools, decision matrix included.
Read articleDeepResearch Output Quality Control: Model Selection, Prompt Tuning, and Citation Verification
DeepResearch output quality varies. A systematic QC path: model selection (Plus vs Pro vs o-series), the 5-part research prompt, citation verification checklist, and the iteration feedback loop.
Read articleDeep Research Prompt Patterns and Quota Strategy: High-Quality Research Reports
Deep Research is ChatGPT's multi-turn research Agent. A 5-part research prompt template, citation verification methods, Plus/Pro quota strategy, and 4 typical research scenarios.
Read articleSubscribe to GPTMap Weekly
One email every Monday: curated OpenAI updates, deep dives, and best practices. No ads, unsubscribe anytime.