GPTMap

How to Read the Benchmark Charts in a Model Launch: Effort Curves, Log Cost Axes, and Safeguard Interventions

Model-launch accuracy-vs-cost charts carry hidden methodology choices. Worked on the 2026-09-01 Claude Fable 5.1 announcement: effort curves, log axes, error bars, scoring conventions, missing cells, safeguard zero-scores.

TL;DR
Launch accuracy-vs-cost charts are evidence with built-in methodology choices. Eight points: lock the effort tier; convert log cost axes; check error bars and trials; partial vs strict; missing cells are unreported; safeguard-intervention zeros; in-house vs leaderboard gaps stay within noise; methodology notes decide comparability. Worked on the 2026-09-01 Claude Fable 5.1 announcement.
Benchmark-chart literacy is the practice of reading the accuracy-vs-cost charts and comparison tables in a model launch as evidence shaped by methodology choices: lock the effort tier and scoring convention first, then verify error bars, trial counts, task-set versions, and safeguard-intervention rules, and only then compare numbers. Demonstrated throughout on the 2026-09-01 Claude Fable 5.1 / Mythos 5.1 launch announcement.

How to

  1. Identify models and tiers

    Map each curve and point to a model and an effort tier; lock the models you compare onto the same tier before reading numbers.

  2. Convert the log cost axis

    On 'log scale', convert visual distance into multiples: read the axis ends (e.g., $0.20-4 or $10-50) and estimate true multiples between points.

  3. Check error bars and trials

    Find the standard error and trial count in the footnotes (e.g., ±3.5-4.5 pts, 3 trials/task); ranking gaps inside the error range carry no conclusion.

  4. Confirm the scoring convention

    Distinguish partial vs strict and no-tools vs with-tools; cite the convention together with the score.

  5. Handle missing cells and zero scores

    Dashes mean unreported, not losing; check whether safeguards intervened and score zero, and read raw capability separately from the safeguarded product.

  6. Separate in-house setups from public leaderboards

    Publisher numbers and leaderboard numbers can legitimately differ; confirm harness and task-set version before choosing which to cite.

  7. Read the methodology notes

    Finish with the footnotes (task-set date, harness, reproduction conditions) and decide which external data this table can be compared against directly.

Model launches now ship with a standard chart type: accuracy-vs-cost scatter curves plus a multi-model comparison table. These charts are evidence -- but they embed a chain of methodology choices: effort tiers, scoring conventions, task-set versions, safeguard rules, error margins. The same model can move by tens of points across conventions. This article uses the 2026-09-01 Claude Fable 5.1 / Mythos 5.1 launch announcement as a worked example and unpacks eight reading points. All numbers are relayed from that announcement.

1. First: Effort Tiers Turn One Model into a Curve

Modern reasoning models have effort tiers -- in this announcement, low / med / high / xhigh / max, with official defaults: High in Claude Code, Medium in Claude Cowork and on Claude.ai. That means "a model" on an accuracy-vs-cost chart is a curve of tier points, not a single point.

Two immediate consequences:

  1. Cross-model comparisons must lock the same tier. The announcement's own claims are tier-scoped -- "at Low or Medium effort, Fable 5.1 achieves results similar to or better than Fable 5's at a much lower cost" says nothing about xhigh/max.
  2. Cross-vendor comparisons must also align tier semantics. GPT-5.6's reasoning.effort and this announcement's effort are both multi-tier designs, but each tier maps to different cost-quality trade-offs; read both vendors' tier documentation before comparing.

2. Log Cost Axes: Visual Distance Is Not Money Distance

Every cost x-axis in the announcement is labeled "USD, log scale" -- and the same page carries two ranges: $10-50 (Terminal-Bench-Science) and $0.20-4 (Humanity's Last Exam). On a log axis, equal spacing means equal multiples, not equal dollars; the visual midpoint of a range is its geometric mean.

Practice: before reading any point, read the axis's two ends and convert positions into multiples.

Three steps for reading the axis (HLE chart example):
1. Axis range: $0.20 -- $4 (log scale)
2. Visual midpoint ≈ geometric mean of $0.20 and $4 ≈ $0.89 -- not $2.1
3. Half the visual distance ≈ a multiple difference (read the ticks), not half the money

3. Error Bars and Trial Counts Are Hard Constraints

The announcement's footnote states Terminal-Bench-Science 0.1 carries a standard error of ±3.5-4.5 points per model, and the corresponding public leaderboard runs 3 trials/task (Claude Code harness). A ranking gap under 2 points therefore supports no conclusion.

The same page provides an official two-setups comparison: the public leaderboard reports Opus 5 at 30.0% and Fable 5 at 21.4%; the publisher's own setup reproduces 29.0% and 24.7% -- within noise, per the footnote. Note Fable 5's two numbers differ by 3.3 points: different harnesses and trial counts genuinely produce different numbers. Before citing any score, ask three questions: which task-set version, which harness, how many trials.

4. Scoring Conventions: partial vs strict Can Differ by 36 Points

OSWorld 2.0 (computer use) appears in the announcement's table under both conventions:

Modelpartialstrict
Fable 5.177.9%41.7%
Fable 572.9%36.1%
Opus 575.4%39.6%
GPT-5.6 Sol

Same model, same task set, and the partial convention scores 36.2 points higher (calculated from the table). Partial typically credits partial completion; strict requires full success. Citing a computer-use score without its convention is an invalid citation. The paired convention also applies to HLE's no-tools / with-tools rows (Fable 5.1: 60.9% / 65.0%).

5. A Missing Cell (—) Is Unreported, Not Losing

The table has no GPT-5.6 Sol numbers for OSWorld 2.0 or Humanity's Last Exam -- the cells are dashes. A missing cell only means the publisher did not run or report that benchmark; reading it as "scored zero" or "avoided it" goes beyond the evidence. The correct move: cite an independent source for that benchmark, or state explicitly that the cell has no public data.

6. Tasks Where Safeguards Intervene Score Zero

The most commonly missed rule: the announcement states Fable 5.1 was evaluated with production safeguards enabled, and tasks where safeguards intervened scored zero -- on OSWorld 2.0 for both Fable tiers and on AutomationBench for Fable 5 -- which Anthropic acknowledges likely reduced those scores.

The same logic explains the two-tier spread: on Terminal-Bench 4.0, Mythos 5.1 (the stronger-safeguard tier) scores 60.9% vs Fable 5.1's 55.8%, officially attributed to earlier, less precise cyber safeguards, with the gap expected to shrink as safeguards improve. Keep "raw model capability" and "the safeguarded product" in separate reading lanes -- cite the one you would actually deploy.

7. Methodology Notes Decide Comparability

The announcement's OSWorld 2.0 footnote: scores are on the benchmark authors' August 2026 task release, with Fable 5 and Opus 5 re-run under the same conditions, and because the task files differ from earlier releases, these numbers are not directly comparable to previous ones. In other words, "77.9% on OSWorld 2.0" holds only within that task-set version; do not mix it with older OSWorld scores circulating online into one timeline table.

Convention checklist (run before citing any score):
[ ] Task-set / dataset version and date
[ ] Scoring convention (partial/strict, no-tools/with-tools)
[ ] Effort tier and default
[ ] Safeguards on? How are interventions scored?
[ ] Error margins and trial counts
[ ] Harness (evaluation framework) name

8. In Practice: Compressing the Table into Three Citable Sentences

Applying the eight points to this announcement's table yields three citable sentences:

# Two common quantities computed directly from the table:
# the partial-strict spread, and cost multiples at a locked tier
fable51_partial, fable51_strict = 77.9, 41.7
print(f"OSWorld 2.0 partial-strict spread: {fable51_partial - fable51_strict:.1f} pts")

# Cost multiple on a log axis: HLE chart range $0.20-$4, point A $0.5, point B $2.0
print(f"Same-tier cost multiple: {2.0 / 0.5:.1f}x (read the ticks, not the visual distance)")

Sentence one: under the official setup, Fable 5.1 leads on the five benchmarks with Sol scores, OSWorld and HLE report no Sol scores, and its results are with production safeguards enabled. Sentence two: every benchmark carries explicit error and convention limits (±3.5-4.5 pts, partial/strict, the 2026-08 task set); align conventions before mixing tables. Sentence three: effort defaults differ by product (Claude Code = High, Claude Cowork / Claude.ai = Medium); lock the tier for cross-vendor pricing. None of the three goes beyond the announcement plus arithmetic -- that is what "reading as evidence" means.

9. Common Mistakes and Troubleshooting

  • Treating a low-tier point as "the model got stronger": cheaper tiers trade score for cost; cross-tier comparison is meaningless.
  • Eyeballing multiples on a log axis: the visual midpoint approximates the geometric mean; eyeballing overestimates.
  • Citing a partial score against someone else's strict score: a 36.2-point convention spread in this example.
  • Reading a missing cell as a loss: a dash is "unreported".
  • Ignoring safeguard-intervention zeros: the safeguarded product's score and the raw model's capability are two different quantities.
  • Mixing task-set versions into one timeline: the methodology notes explicitly forbid it.

10. Next Steps

Key points

  • Effort tiers (low / med / high / xhigh / max) turn 'a model' into a cost-quality curve: cross-vendor price comparisons must lock the same tier or they do not hold
  • Cost axes are almost always log scale: the same announcement page carries $0.20-4 and $10-50 x-axes, so a visual midpoint can hide a 5-10x gap
  • Error bars are a hard constraint: Terminal-Bench-Science 0.1 carries ±3.5-4.5 pts standard error per model and the public leaderboard runs 3 trials/task
  • Partial and strict are two different scores: OSWorld 2.0 shows Fable 5.1 at 77.9% partial vs 41.7% strict -- a 36.2-point spread; always cite the convention
  • A missing cell (dash) means unreported, not losing: the announcement table has no GPT-5.6 Sol scores for OSWorld 2.0 or HLE
  • Tasks where safety safeguards intervene score zero: Fable 5.1 ran with production safeguards, which Anthropic says likely depressed its scores; the Mythos-vs-Fable gap reflects safeguard granularity
  • An in-house setup and the public leaderboard can differ: the leaderboard reports Opus 5 at 30.0% / Fable 5 at 21.4%, Anthropic's setup reproduces 29.0% / 24.7%, all within noise -- ask which harness and how many trials
  • Methodology notes decide comparability: OSWorld 2.0 used the benchmark authors' August 2026 task release and is officially not directly comparable to earlier versions

Frequently asked questions

Because reasoning effort tiers (low / med / high / xhigh / max) unfold a single model into a cost-quality curve: higher effort means higher scores and higher cost. Step one of chart reading is identifying which curve belongs to which model and which point to which tier. Cross-model price comparisons must lock the same tier -- a claim like 'low-effort Fable 5.1 matches Fable 5 for less cost' holds only on those tiers and may not extrapolate to xhigh/max.

Official references

Related articles

Subscribe to GPTMap Weekly

One email every Monday: curated OpenAI updates, deep dives, and best practices. No ads, unsubscribe anytime.

GPTMap EditorialPublished 2026-09-02 7 min read
Test environment (EEAT)
Last tested: 2026-09-02
Model used: gpt-5.6