Repository navigation
fix(pricing): add Haiku 5.5, Sonnet 5.5 and GPT-6.1 Sol, re-verify every row - #220
Conversation
…ery row Haiku 5.5 (released 2026-10-07) is priced at its up-to-100K-prompt tier; prompts over 100K bill 5x and read low, like the other long-context tiers the card does not model. Sonnet 5.5's cache read is the 0.05x rate that took effect with the Haiku 5.5 launch. GPT-6.1 Sol is new on OpenAI's card. Every other row still matches its public list price. Co-Authored-By: Claude Opus 5.5 <[email protected]>
|
Claude finished @bai-uipath's task in 1m 11s —— View job Code Review in Progress
|
uipreliga
left a comment
There was a problem hiding this comment.
Review: coder_eval — pr:220 (2 files) axis:1,2,3,4,5,6,7,8
Scope: pr:220 (2 files) axis:1,2,3,4,5,6,7,8 · branch bai/pricing-haiku-5-5 · 2902fa0 · 2026-10-07T18:56Z · workflow variant
Change class: trivial — adds three rate rows to the pricing table (and their generated mirror) and bumps a docstring date; no code path changes
The codebase is healthy: seven of eight axes score 10/10 and there are no critical, high or medium findings, so no finding can change a task's score or final_status for identical agent output. The only real risk is that Haiku 5.5 is priced at its flat <=100K-prompt rate, so LiteLLM-repriced, backfilled, judge and simulator costs, and the max_usd gate that uses them, can read up to 5x low while reporting cost_complete=true. The other two findings are comment nits in src/coder_eval/pricing.py. Bottom line: ship it, and add a short follow-up on tiered-model cost honesty.
Summary
| Axis | Score | 🔴 | 🟠 | 🟡 | 🔵 | Top Issue |
|---|---|---|---|---|---|---|
| 1. Code Quality & Style | 9.8 / 10 | 0 | 0 | 0 | 2 | gpt-6.1-sol row has a 0.05x cache-read ratio with no annotation, and it sits under gpt-6-astra's CAVEAT comment |
| 2. Type Safety | 10 / 10 | 0 | 0 | 0 | 0 | — |
| 3. Test Health | 10 / 10 | 0 | 0 | 0 | 0 | — |
| 4. Security | 10 / 10 | 0 | 0 | 0 | 0 | — |
| 5. Architecture & Design | 10 / 10 | 0 | 0 | 0 | 0 | — |
| 6. Error Handling & Resilience | 10 / 10 | 0 | 0 | 0 | 0 | — |
| 7. API Surface & Maintainability | 10 / 10 | 0 | 0 | 0 | 0 | — |
| 8. Evaluation Harness Quality | 9.9 / 10 | 0 | 0 | 0 | 1 | Haiku 5.5 is priced only at its <=100K-prompt tier, and the 5x long-context tier lives only in a Python comment, so board costs and the max_usd gate read up to 5x low |
Overall Score: 10 / 10 · Weakest Axis: Code Quality & Style at 9.8 / 10
Totals: 🔴 0 · 🟠 0 · 🟡 0 · 🔵 3 across 8 axes.
Blockers
None.
Non-blocking, but please consider before merge
None.
Nits
- [Axis 1] gpt-6.1-sol row has a 0.05x cache-read ratio with no annotation, and it sits under gpt-6-astra's CAVEAT comment (
src/coder_eval/pricing.py:126) — Line 126"gpt-6.1-sol": ModelPricing(2.0, 10.0, 2.50, 0.10),prices cache hits at 0.05x input (0.10 / 2.0). Its sibling on line 127 ("gpt-6-sol": ModelPricing(2.0, 10.0, 2.50, 0.20)) and the rest of the OpenAI block use 0.1x. This file marks every off-norm cache rate with a one-line guard so that a later edit does not 'correct' it. Examples: line 67# Opus 5.5 prices cache hits at 0.05x input, not the usual 0.1x.and line 105# The one OpenAI entry whose cached rate is 25% of input, not 10%.The new row has no such guard. It also comes directly after line 124# CAVEAT: flat rate; GPT-6's >272K-input tier (2x input / 1.5x output) is not modelled and reads low.and before gpt-6-sol, so a reader cannot tell if the astra caveat also applies to 6.1-sol. Add a short guard comment above line 126 (for example# 6.1 Sol prices cache hits at 0.05x input, not 6 Sol's 0.1x.). If the >272K tier also applies, say so in that comment. If it does not, move the row so that it is not under the astra caveat. Cosmetic: the rate itself is correct and not disputed. - [Axis 1] Sonnet 5.5 comment records a dated price-change history and does not use the HAZARD convention (
src/coder_eval/pricing.py:80) — Line 80# $2/$10, NOT Sonnet 4.6's $3/$15. 5.5's cache hits fell to 0.05x input on 2026-10-07.describes a change over time. The row it annotates (line 81,"claude-sonnet-5-5": ModelPricing(2.0, 10.0, 2.50, 0.10)) is new in this PR, so the earlier rate implied by 'fell' does not appear anywhere in the file. CLAUDE.md says a comment states the contract, not the history. If the date matters because runs before 2026-10-07 are now re-priced at the lower rate, the file already has a convention for that: the# HAZARD: a single CURRENT-rate card with no effective date ...comment on lines 118-120. Otherwise, write the comment as a rule, like line 67:# $2/$10, NOT Sonnet 4.6's $3/$15. Sonnet 5.5 prices cache hits at 0.05x input, not 5's 0.1x.This is a comment-quality nit with no effect on how the code behaves. - [Axis 8] Haiku 5.5 is priced only at its <=100K-prompt tier, and the 5x long-context tier lives only in a Python comment, so board costs and the max_usd gate read up to 5x low (
src/coder_eval/pricing.py:87-88) — Lines 87-88 at pr-220:# CAVEAT: the <=100K-prompt tier; a prompt over 100K bills every bucket at 5x, so long contexts read low.and"claude-haiku-5-5": ModelPricing(0.10, 0.50, 0.125, 0.01),. The caveat is accurate, andcalculate_costcannot model the tier, because it receives buckets summed per turn, not per request. But the consequence for the harness is larger than the earlier flat-rate caveats (gpt-5.5 and gpt-6-astra at 2x past 272K). Coding-agent prompts often go past 100K, and here every bucket bills at 5x. Before this PR the model was unpriced, so cost was None, themax_usdgate was skipped with a warning (orchestrator.py:1169-1178), andcost_completewas false, which is an honest result. After this PR, the paths that price from the rate card book a figure that can be up to 5x low, and the row reports cost_complete=true. Those paths are Claude Code behind LiteLLM (_reprice_for_litellm, claude_code_agent.py:1513, called at :615), the timeout/kill backfill (_backfill_cost, :1494), judge usage (evaluation/judge_usage.py:59) and simulator cost. In that stateBudgetExceededError("usd", ...)can fail to fire when the real bill is over the cap. The Claude Code direct-Anthropic path uses the SDK's owntotal_cost_usd, so it is not affected. This does not change score or final_status for identical agent output, so it is below the Axis 8 blocker bar. Options: (a) document in docs/TASK_DEFINITION_GUIDE.md § Run Limits thatmax_usdis a lower bound for tiered models, or (b) add atiered/flat_rate_floormarker toModelPricing. With (b),check_pricing_coveragecan warn andcost_completecan report false when such a model is used withmax_usd. The current state is acceptable for a pricing-only PR. Nightly impact: no change to the schema, task.json, CLI or Docker contract. Runs on these three models move from cost=null to a numeric cost. The image version preflight already forces a lockstep image rebuild at release.
What's Missing
Parallel paths:
- 🟡 The new claude-haiku-5-5 row does not reach the Claude Code direct-Anthropic path, which is the most common harness. That path takes the SDK's
model_usage.costUSD/total_cost_usdas-is (claude_code_agent.py:1440-1443) and uses the rate card only through_backfill_costwhen no cost is reported. The pinned CLI (CLAUDE_CODE_VERSION=2.1.281in docker/Dockerfile:33 and docker/Dockerfile.runtime:69,112) does not recognise Haiku 5.5 and bills it at Opus 5.5 rates, about 40x list. As a result, Haiku 5.5 rows on Claude Code record cost about 40x too high in run.json and on the board.max_usdalso fires early with COST_BUDGET_EXCEEDED, which changes final_status on runs that are far under budget. The PR body says this is out of scope, but nothing in the code or docs guards it. Do one of these: bump CLAUDE_CODE_VERSION in lockstep, or reprice from buckets (as_reprice_for_litellmdoes) when the SDK reportscostBasis: unknown. At minimum, add a known-issue note to docs/agents/CLAUDE_CODE.md. (trigger: src/coder_eval/pricing.py) - 🔵 Python and the board look up model ids differently. evalboard/lib/pricing.ts
resolvePricingstrips a trailing -YYYYMMDD before the lookup. Python_normalize_modeldoes not, which is why Haiku 4.5 needs a separate dated twin row (claude-haiku-4-5-20251001). The three new rows have no dated twin. If the Anthropic API or the CLI reports Haiku 5.5 or Sonnet 5.5 under a dated snapshot id (as Haiku 4.5 does), the board prices the run, but Python returns cost None and skipsmax_usd. Either confirm that the new models have no dated id, or make_normalize_modelstrip the date suffix the way the board does, so the dated twin rows can be deleted. (trigger: src/coder_eval/pricing.py)
Downstream consumers:
- 🔵 These three models were unpriced before, so their cost was None and the
max_usdgate was skipped with a warning. Now they are priced, so the gate is armed on every path that prices from the rate card: codex, antigravity, delegate, Claude Code behind LiteLLM, the timeout/kill backfill, judge and simulator. For Haiku 5.5, past 100K, that cost is only a lower bound. docs/TASK_DEFINITION_GUIDE.md § Run Limits (lines 270 and 295-296) saysmax_usdneeds per-turn cost, but it does not say the gate under-counts for tiered models (Haiku 5.5, GPT-5.5, GPT-6 astra, Gemini 3.1 Pro). Add that note in the same change. (trigger: src/coder_eval/pricing.py) (restates: Axis 8: Haiku 5.5 is priced only at its <=100K-prompt tier) - 🔵 Before this PR,
resolvePricingdid not know claude-sonnet-5-5, so the board showed '—' for Sonnet 5.5 runs. It now prices every Sonnet 5.5 run with a single current rate, including runs from before 2026-10-07. ThroughmessageCostUsd,tokenBucketUsdand the thinking-cost simulator, those older runs show cache-read cost at $0.10, not the $0.20 that was actually billed. The board gives no effective-date marker. The rate card has no effective-date concept, and this drift is not flagged with the file's HAZARD convention, so a consumer of evalboard/lib/pricing.ts cannot tell that historical Sonnet 5.5 costs read low. (trigger: evalboard/lib/pricing.generated.ts) (restates: Axis 1: Sonnet 5.5 comment records a dated price-change history and does not use the HAZARD convention)
Tests:
- 🔵 No test pins the deliberate off-norm cache-read ratios. Sonnet 5.5 and GPT-6.1 Sol are new rows at 0.05x input; Opus 5.5 (0.05x), Fable 5.1 (0.025x) and the 25% OpenAI entry already exist. The only thing stopping a later edit from 'correcting' them to the usual 0.1x is a comment, and GPT-6.1 Sol does not have one. CE065 compares only the mirror with pricing.py, not the values themselves, so a 'correction' passes every gate. A small table-driven test in tests/test_pricing_registry.py that asserts the expected off-norm ratio for each listed id would catch this. (trigger: src/coder_eval/pricing.py) (restates: Axis 1: gpt-6.1-sol row has a 0.05x cache-read ratio with no annotation)
Nightly pipeline:
- 🔵 The PR description does not say what happens to the nightly run. For nightly rows on these three models,
costchanges from null to a number andcost_completechanges from false to true, which is a visible change in run.json and the board totals. Any nightly task that setsmax_usdis now gated where before it was skipped, so its final_status can change to COST_BUDGET_EXCEEDED. Claude Code rows on Haiku 5.5 will also book about 40x cost until the CLI is bumped. The schema, the task.json shape and the CLI stay the same. The pricing change ships only with a lockstep image rebuild at release. Say this in the PR so the coder-eval-uipath nightly owners can expect the change in cost columns. (trigger: src/coder_eval/pricing.py) (restates: Axis 8: Haiku 5.5 is priced only at its <=100K-prompt tier)
Harness & Lint Improvements
Static checks (lint / type):
- [ce-lint] CE069 (next free number; CE067 is unused but skip it to avoid confusion). Add a whole-tree
@pytest.mark.lintclassTestCE069OffNormCacheReadDeclaredin tests/test_custom_lint.py, with its tableOFF_NORM_CACHE_READ: dict[str, str](key -> reason) besideDELIBERATELY_UNMIRROREDin tests/lint/pricing_mirror.py. Usebuiltin_rates()and leave outper_request_billingrows andDELEGATE_MODEL_IDS(that block has its own notes pointer). Every row whose cache_read_per_mtok / input_per_mtok != 0.1 must be a key of the table. The reason must contain the ratio literal (for example '0.05x'), so that it states a rule and not a history. Every key must still be priced and still off-norm, using the same stale-entry check as_assert_exemptions_are_live. Then delete the free-text ratio guard comments that the table replaces (pricing.py:63-64 Fable, :67 Opus 5.5, :105 codex-mini, :108 Pro tiers). A run on pr-220 gives these off-norm, non-exempt rows: claude-fable-5-1, claude-opus-5-5, claude-sonnet-5-5, claude-3-haiku-20240307 (0.12x, which has no annotation before this PR either), codex-mini-latest, gpt-5.5-pro, gpt-5.4-pro, gpt-6.1-sol, the three Bedrock rows and the two jev rows. Prevents: Finding 1: gpt-6.1-sol at pricing.py:126 has a 0.05x cache-read ratio with no annotation. The rule fails until the row is declared. Finding 2, in part: the Sonnet 5.5 row at pricing.py:81 must get a reason that contains '0.05x', which puts the ratio as a rule and not as a dated change. It also finds the old claude-3-haiku-20240307 0.12x row, which has no annotation. - [ce-lint] CE070: put tiered pricing in data, not in comments. Add
tier_threshold_tokens: int | None = None(orflat_rate_floor: bool = False) toModelPricingin src/coder_eval/pricing.py, and set it on every rate that has a long-context tier. Add a lint class in tests/test_custom_lint.py that fails on any own-line comment in src/coder_eval/pricing.py that matches#.*\b(tier|surcharge|long-context)\b|>\s*\d+K. After that, the only place a tier can be recorded is the field. Today the regex matches gpt-5.5 (:110), gpt-6-astra (:124), the Gemini Pro block (:127) and, on pr-220, claude-haiku-5-5 (:87). Delete those CAVEAT comments when you set the field. Also show the field in the evalboard mirror (render_pricing) so that the board can mark such a cost as a floor. Prevents: Finding 3: Haiku 5.5's 5x tier above 100K tokens is recorded only in a Python comment at pricing.py:87-88, so no code can act on it. It also removes the scoping problem in Finding 1: gpt-6.1-sol sits under gpt-6-astra's group CAVEAT comment, and no group comment can be read the wrong way when the tier is a per-row field. - [ci-gate] Extend
make docs-budget(tests/lint/prose_budget.py) with a 'no dated history' rule. An own-line comment under src/ that has an ISO date (\b20\d{2}-\d{2}-\d{2}\b) must also have an expiry or contract marker (through|until|re-check|expires|HAZARD). There are no false positives on main today: the only dated comments, pricing.py:117 ('promotional through at least 2026-11-21; re-check then') and pricing.py:128 ('discounted by half through 2026-12-31'), are expiry markers. Known blind spot: history without a date ('used to', 'was') needs semantic judgment and stays a review item. Prevents: Finding 2: pricing.py:80 '5.5's cache hits fell to 0.05x input on 2026-10-07' is a dated change history, and CLAUDE.md says a comment states the contract, not the history. The new rule fails on this line.
Harness improvements (not statically reachable):
- When a model has the tier field from CE070 and its cost comes from the rate card (
_reprice_for_litellmclaude_code_agent.py:1513,_backfill_cost:1494, evaluation/judge_usage.py:59, the simulator cost path), setcost_complete=falsein run_record.py_cost_complete. Also emit the existing one-shotmax_usdwarning (orchestrator.py:1169-1178) that says the budget check is a lower bound. Add a unit test: put a turn with a prompt over 100K tokens on claude-haiku-5-5 through the LiteLLM reprice path, then check that cost_complete is False and that the warning is emitted. Why not static: Whether a prompt goes over the tier boundary is known only at runtime.calculate_costalso gets token buckets summed per turn, not per request, so no AST or type check can tell if a priced figure is a floor. Prevents: Finding 3: before this PR the model had no price, cost was None and cost_complete was false, which is honest. After this PR, a cost up to 5x too low is booked with cost_complete=true, andBudgetExceededError("usd")can fail to fire when the real bill is over the cap. - Generate a short list of tiered models (models whose
max_usdvalue is a floor) from the CE070 field into docs/TASK_DEFINITION_GUIDE.md § Run Limits. Use the same generation and drift-check method asmake plugin-reference/make pricing-mirror, so the list cannot fall out of sync with pricing.py. Why not static: Static checks can stop drift in the generated list. But the decision to tell users that max_usd is a lower bound is about wording and documentation, and only a review can confirm it. A static check cannot find a missing warning in prose until the list is generated. Prevents: Finding 3: today an author who setsmax_usdon claude-haiku-5-5, gpt-5.5, gpt-6-astra or Gemini Pro is not told that the cap can read up to 5x low (2x for the GPT/Gemini tiers).
Top 5 Priority Actions
- Document in docs/TASK_DEFINITION_GUIDE.md under Run Limits that max_usd is only a lower bound for tiered models. Haiku 5.5 is the main case: at src/coder_eval/pricing.py:87-88 it is priced at its <=100K-prompt rate, but past 100K every bucket bills at 5x, so BudgetExceededError("usd") can fail to fire on the LiteLLM reprice path (claude_code_agent.py:1513), the backfill path (:1494) and the judge path (evaluation/judge_usage.py:59).
- Add a
tiered(flat-rate floor) marker to ModelPricing and set it on claude-haiku-5-5, gpt-5.5 and gpt-6-astra, so that check_pricing_coverage warns and cost_complete reports false when one of these models runs under max_usd. Today these rows produce a confident number that can be too low, where before they produced an honest null. - Add a guard comment above src/coder_eval/pricing.py:126 that marks the off-norm rate, for example
# 6.1 Sol prices cache hits at 0.05x input, not 6 Sol's 0.1x.. Also say whether the >272K-tier caveat on line 124 applies to gpt-6.1-sol, or move the row out from under that caveat. - Change the Sonnet 5.5 comment at src/coder_eval/pricing.py:80 from a dated history (
fell to 0.05x ... on 2026-10-07) to a rule in the style of line 67. If the effective date matters for re-pricing older runs, use the existing HAZARD convention (lines 118-120) instead. - After the pricing.py edits, run
make pricing-mirrorandmake verifyso that CE065 confirms evalboard/lib/pricing.generated.ts matches the three new rows. Also tell nightly consumers that runs on these models now report a numeric cost instead of cost=null.
Stats: 0 🔴 · 0 🟠 · 0 🟡 · 3 🔵 across 8 axes reviewed.
uipreliga
left a comment
There was a problem hiding this comment.
fix what you agree with and 🚢

Claude Haiku 5.5 launched today (2026-10-07) and Claude Sonnet 5.5 and GPT-6.1 Sol have no rows yet, so any tier-3 harness (codex, antigravity, delegate) and every llm_judge / simulator overhead on them books
cost: None. This adds the three and re-checks every other row against the providers' pages.Changes
claude-haiku-5-5: $0.10 in / $0.50 out / $0.125 cache write / $0.01 cache read. That is the tier for prompts up to 100K tokens. A request over 100K bills every bucket at 5x ($0.50 / $2.50 / $0.625 / $0.05). A flat row can't express that, so long-context Haiku 5.5 runs read low. This is the same caveat the card already carries for GPT-5.5, GPT-6 and Gemini 3.1 Pro.claude-sonnet-5-5: $2 / $10 / $2.50 / $0.10. Anthropic cut Sonnet 5.5 cache reads from $0.20 to $0.10 (0.05x input) with the Haiku 5.5 launch. The card has no effective date, so a Sonnet 5.5 run from before today re-prices slightly low.gpt-6.1-sol: $2 / $10 / $2.50 / $0.10, from OpenAI's standard tier.pricing.generated.tsregenerated withmake pricing-mirror.Re-verified, unchanged
/api/v1/modelsheadline now follows the cheapest promotional endpoint (GLM 5.2 $0.171 / $7.20, Kimi K3 $0.50 / $15). Those three rows are only themax_usdpre-flight fallback behindper_request_billing, and a promo-low figure would weaken that guard. Real cost still comes per call.Validation
tests/test_pricing_mirror.py,tests/test_pricing_registry.pytests/test_custom_lint.py(incl. CE065 mirror drift)pytest -k "pric or cost"prose_budget, ruff check / formateu.andglobal.anthropic.claude-haiku-5-5return 200;us.400Not addressed here: Claude Code 2.1.281 (the pinned CLI) doesn't recognize Haiku 5.5 yet. It books it at Opus 5.5 rates (
costBasis: unknown, about 40x list), so claude-code runs on Haiku 5.5 need a CLI bump or a reprice from tokens.Sources: Anthropic pricing, Haiku 5.5 overview, OpenAI pricing, Gemini pricing, OpenRouter models
🤖 Generated with Claude Code