Skip to content

fix(pricing): add Haiku 5.5, Sonnet 5.5 and GPT-6.1 Sol, re-verify every row - #220

Merged
bai-uipath merged 1 commit into
mainfrom
bai/pricing-haiku-5-5
Oct 7, 2026
Merged

bai-uipath merged 1 commit into
mainfrom
bai/pricing-haiku-5-5

Conversation

@bai-uipath

Copy link
Copy Markdown
Collaborator

Claude Haiku 5.5 launched today (2026-10-07) and Claude Sonnet 5.5 and GPT-6.1 Sol have no rows yet, so any tier-3 harness (codex, antigravity, delegate) and every llm_judge / simulator overhead on them books cost: None. This adds the three and re-checks every other row against the providers' pages.

Changes

  • New claude-haiku-5-5: $0.10 in / $0.50 out / $0.125 cache write / $0.01 cache read. That is the tier for prompts up to 100K tokens. A request over 100K bills every bucket at 5x ($0.50 / $2.50 / $0.625 / $0.05). A flat row can't express that, so long-context Haiku 5.5 runs read low. This is the same caveat the card already carries for GPT-5.5, GPT-6 and Gemini 3.1 Pro.
  • New claude-sonnet-5-5: $2 / $10 / $2.50 / $0.10. Anthropic cut Sonnet 5.5 cache reads from $0.20 to $0.10 (0.05x input) with the Haiku 5.5 launch. The card has no effective date, so a Sonnet 5.5 run from before today re-prices slightly low.
  • New gpt-6.1-sol: $2 / $10 / $2.50 / $0.10, from OpenAI's standard tier.
  • Docstring re-verification date bumped to 2026-10-07. pricing.generated.ts regenerated with make pricing-mirror.

Re-verified, unchanged

Block Result (2026-10-07)
Anthropic (Fable, Opus, Sonnet 5 / 4.x, Haiku 4.5 / 3.5, legacy) every row matches
OpenAI GPT-6 / 5.6 / 5.5 / 5.4 / 5 / 5.3-codex every row matches; 5.6 Sol promo still runs through at least 2026-11-21
OpenAI 5.2 / 5.1 codex, codex-mini gone from the page; last published rate kept
Gemini every list rate matches; the 3.6 to 3.8 Flash 50% discount still ends 2026-12-31
OpenRouter kimi-k3 / glm-5.2 / deepseek-v4-pro kept as is (see below)
Delegate block, Bedrock open-weight not publicly listed; kept as is
  • OpenRouter rows left alone on purpose: the /api/v1/models headline now follows the cheapest promotional endpoint (GLM 5.2 $0.171 / $7.20, Kimi K3 $0.50 / $15). Those three rows are only the max_usd pre-flight fallback behind per_request_billing, and a promo-low figure would weaken that guard. Real cost still comes per call.
  • Pricing-page inconsistency: Anthropic's model table still lists Sonnet 5.5 cache hits at $0.20, but the prompt-caching section on the same page and the launch coverage both give $0.10. This PR uses $0.10.

Validation

Check Result
tests/test_pricing_mirror.py, tests/test_pricing_registry.py 30 passed
tests/test_custom_lint.py (incl. CE065 mirror drift) 735 passed
pytest -k "pric or cost" 190 passed
prose_budget, ruff check / format clean
Bedrock eu-north-1 probe eu. and global.anthropic.claude-haiku-5-5 return 200; us. 400

Not addressed here: Claude Code 2.1.281 (the pinned CLI) doesn't recognize Haiku 5.5 yet. It books it at Opus 5.5 rates (costBasis: unknown, about 40x list), so claude-code runs on Haiku 5.5 need a CLI bump or a reprice from tokens.

Sources: Anthropic pricing, Haiku 5.5 overview, OpenAI pricing, Gemini pricing, OpenRouter models

🤖 Generated with Claude Code

…ery row

Haiku 5.5 (released 2026-10-07) is priced at its up-to-100K-prompt tier;
prompts over 100K bill 5x and read low, like the other long-context tiers
the card does not model. Sonnet 5.5's cache read is the 0.05x rate that
took effect with the Haiku 5.5 launch. GPT-6.1 Sol is new on OpenAI's card.
Every other row still matches its public list price.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@bai-uipath
bai-uipath marked this pull request as ready for review October 7, 2026 18:40
@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Claude finished @bai-uipath's task in 1m 11s —— View job


Code Review in Progress

Todo List:

  • Read .github/code_review.md for review standards
  • Read CLAUDE.md for project conventions
  • Review full diff with git diff origin/main...HEAD
  • Read full context of changed files
  • Verify pricing accuracy against provider documentation
  • Check cross-file consistency
  • Analyze "what's missing"
  • Post comprehensive review

Starting review...

@uipreliga uipreliga left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: coder_eval — pr:220 (2 files) axis:1,2,3,4,5,6,7,8

Scope: pr:220 (2 files) axis:1,2,3,4,5,6,7,8 · branch bai/pricing-haiku-5-5 · 2902fa0 · 2026-10-07T18:56Z · workflow variant

Change class: trivial — adds three rate rows to the pricing table (and their generated mirror) and bumps a docstring date; no code path changes

The codebase is healthy: seven of eight axes score 10/10 and there are no critical, high or medium findings, so no finding can change a task's score or final_status for identical agent output. The only real risk is that Haiku 5.5 is priced at its flat <=100K-prompt rate, so LiteLLM-repriced, backfilled, judge and simulator costs, and the max_usd gate that uses them, can read up to 5x low while reporting cost_complete=true. The other two findings are comment nits in src/coder_eval/pricing.py. Bottom line: ship it, and add a short follow-up on tiered-model cost honesty.

Summary

Axis Score 🔴 🟠 🟡 🔵 Top Issue
1. Code Quality & Style 9.8 / 10 0 0 0 2 gpt-6.1-sol row has a 0.05x cache-read ratio with no annotation, and it sits under gpt-6-astra's CAVEAT comment
2. Type Safety 10 / 10 0 0 0 0 —
3. Test Health 10 / 10 0 0 0 0 —
4. Security 10 / 10 0 0 0 0 —
5. Architecture & Design 10 / 10 0 0 0 0 —
6. Error Handling & Resilience 10 / 10 0 0 0 0 —
7. API Surface & Maintainability 10 / 10 0 0 0 0 —
8. Evaluation Harness Quality 9.9 / 10 0 0 0 1 Haiku 5.5 is priced only at its <=100K-prompt tier, and the 5x long-context tier lives only in a Python comment, so board costs and the max_usd gate read up to 5x low

Overall Score: 10 / 10 · Weakest Axis: Code Quality & Style at 9.8 / 10
Totals: 🔴 0 · 🟠 0 · 🟡 0 · 🔵 3 across 8 axes.

Blockers

None.

Non-blocking, but please consider before merge

None.

Nits

  1. [Axis 1] gpt-6.1-sol row has a 0.05x cache-read ratio with no annotation, and it sits under gpt-6-astra's CAVEAT comment (src/coder_eval/pricing.py:126) — Line 126 "gpt-6.1-sol": ModelPricing(2.0, 10.0, 2.50, 0.10), prices cache hits at 0.05x input (0.10 / 2.0). Its sibling on line 127 ("gpt-6-sol": ModelPricing(2.0, 10.0, 2.50, 0.20)) and the rest of the OpenAI block use 0.1x. This file marks every off-norm cache rate with a one-line guard so that a later edit does not 'correct' it. Examples: line 67 # Opus 5.5 prices cache hits at 0.05x input, not the usual 0.1x. and line 105 # The one OpenAI entry whose cached rate is 25% of input, not 10%. The new row has no such guard. It also comes directly after line 124 # CAVEAT: flat rate; GPT-6's >272K-input tier (2x input / 1.5x output) is not modelled and reads low. and before gpt-6-sol, so a reader cannot tell if the astra caveat also applies to 6.1-sol. Add a short guard comment above line 126 (for example # 6.1 Sol prices cache hits at 0.05x input, not 6 Sol's 0.1x.). If the >272K tier also applies, say so in that comment. If it does not, move the row so that it is not under the astra caveat. Cosmetic: the rate itself is correct and not disputed.
  2. [Axis 1] Sonnet 5.5 comment records a dated price-change history and does not use the HAZARD convention (src/coder_eval/pricing.py:80) — Line 80 # $2/$10, NOT Sonnet 4.6's $3/$15. 5.5's cache hits fell to 0.05x input on 2026-10-07. describes a change over time. The row it annotates (line 81, "claude-sonnet-5-5": ModelPricing(2.0, 10.0, 2.50, 0.10)) is new in this PR, so the earlier rate implied by 'fell' does not appear anywhere in the file. CLAUDE.md says a comment states the contract, not the history. If the date matters because runs before 2026-10-07 are now re-priced at the lower rate, the file already has a convention for that: the # HAZARD: a single CURRENT-rate card with no effective date ... comment on lines 118-120. Otherwise, write the comment as a rule, like line 67: # $2/$10, NOT Sonnet 4.6's $3/$15. Sonnet 5.5 prices cache hits at 0.05x input, not 5's 0.1x. This is a comment-quality nit with no effect on how the code behaves.
  3. [Axis 8] Haiku 5.5 is priced only at its <=100K-prompt tier, and the 5x long-context tier lives only in a Python comment, so board costs and the max_usd gate read up to 5x low (src/coder_eval/pricing.py:87-88) — Lines 87-88 at pr-220: # CAVEAT: the <=100K-prompt tier; a prompt over 100K bills every bucket at 5x, so long contexts read low. and "claude-haiku-5-5": ModelPricing(0.10, 0.50, 0.125, 0.01),. The caveat is accurate, and calculate_cost cannot model the tier, because it receives buckets summed per turn, not per request. But the consequence for the harness is larger than the earlier flat-rate caveats (gpt-5.5 and gpt-6-astra at 2x past 272K). Coding-agent prompts often go past 100K, and here every bucket bills at 5x. Before this PR the model was unpriced, so cost was None, the max_usd gate was skipped with a warning (orchestrator.py:1169-1178), and cost_complete was false, which is an honest result. After this PR, the paths that price from the rate card book a figure that can be up to 5x low, and the row reports cost_complete=true. Those paths are Claude Code behind LiteLLM (_reprice_for_litellm, claude_code_agent.py:1513, called at :615), the timeout/kill backfill (_backfill_cost, :1494), judge usage (evaluation/judge_usage.py:59) and simulator cost. In that state BudgetExceededError("usd", ...) can fail to fire when the real bill is over the cap. The Claude Code direct-Anthropic path uses the SDK's own total_cost_usd, so it is not affected. This does not change score or final_status for identical agent output, so it is below the Axis 8 blocker bar. Options: (a) document in docs/TASK_DEFINITION_GUIDE.md § Run Limits that max_usd is a lower bound for tiered models, or (b) add a tiered/flat_rate_floor marker to ModelPricing. With (b), check_pricing_coverage can warn and cost_complete can report false when such a model is used with max_usd. The current state is acceptable for a pricing-only PR. Nightly impact: no change to the schema, task.json, CLI or Docker contract. Runs on these three models move from cost=null to a numeric cost. The image version preflight already forces a lockstep image rebuild at release.

What's Missing

Parallel paths:

  • 🟡 The new claude-haiku-5-5 row does not reach the Claude Code direct-Anthropic path, which is the most common harness. That path takes the SDK's model_usage.costUSD / total_cost_usd as-is (claude_code_agent.py:1440-1443) and uses the rate card only through _backfill_cost when no cost is reported. The pinned CLI (CLAUDE_CODE_VERSION=2.1.281 in docker/Dockerfile:33 and docker/Dockerfile.runtime:69,112) does not recognise Haiku 5.5 and bills it at Opus 5.5 rates, about 40x list. As a result, Haiku 5.5 rows on Claude Code record cost about 40x too high in run.json and on the board. max_usd also fires early with COST_BUDGET_EXCEEDED, which changes final_status on runs that are far under budget. The PR body says this is out of scope, but nothing in the code or docs guards it. Do one of these: bump CLAUDE_CODE_VERSION in lockstep, or reprice from buckets (as _reprice_for_litellm does) when the SDK reports costBasis: unknown. At minimum, add a known-issue note to docs/agents/CLAUDE_CODE.md. (trigger: src/coder_eval/pricing.py)
  • 🔵 Python and the board look up model ids differently. evalboard/lib/pricing.ts resolvePricing strips a trailing -YYYYMMDD before the lookup. Python _normalize_model does not, which is why Haiku 4.5 needs a separate dated twin row (claude-haiku-4-5-20251001). The three new rows have no dated twin. If the Anthropic API or the CLI reports Haiku 5.5 or Sonnet 5.5 under a dated snapshot id (as Haiku 4.5 does), the board prices the run, but Python returns cost None and skips max_usd. Either confirm that the new models have no dated id, or make _normalize_model strip the date suffix the way the board does, so the dated twin rows can be deleted. (trigger: src/coder_eval/pricing.py)

Downstream consumers:

  • 🔵 These three models were unpriced before, so their cost was None and the max_usd gate was skipped with a warning. Now they are priced, so the gate is armed on every path that prices from the rate card: codex, antigravity, delegate, Claude Code behind LiteLLM, the timeout/kill backfill, judge and simulator. For Haiku 5.5, past 100K, that cost is only a lower bound. docs/TASK_DEFINITION_GUIDE.md § Run Limits (lines 270 and 295-296) says max_usd needs per-turn cost, but it does not say the gate under-counts for tiered models (Haiku 5.5, GPT-5.5, GPT-6 astra, Gemini 3.1 Pro). Add that note in the same change. (trigger: src/coder_eval/pricing.py) (restates: Axis 8: Haiku 5.5 is priced only at its <=100K-prompt tier)
  • 🔵 Before this PR, resolvePricing did not know claude-sonnet-5-5, so the board showed '—' for Sonnet 5.5 runs. It now prices every Sonnet 5.5 run with a single current rate, including runs from before 2026-10-07. Through messageCostUsd, tokenBucketUsd and the thinking-cost simulator, those older runs show cache-read cost at $0.10, not the $0.20 that was actually billed. The board gives no effective-date marker. The rate card has no effective-date concept, and this drift is not flagged with the file's HAZARD convention, so a consumer of evalboard/lib/pricing.ts cannot tell that historical Sonnet 5.5 costs read low. (trigger: evalboard/lib/pricing.generated.ts) (restates: Axis 1: Sonnet 5.5 comment records a dated price-change history and does not use the HAZARD convention)

Tests:

  • 🔵 No test pins the deliberate off-norm cache-read ratios. Sonnet 5.5 and GPT-6.1 Sol are new rows at 0.05x input; Opus 5.5 (0.05x), Fable 5.1 (0.025x) and the 25% OpenAI entry already exist. The only thing stopping a later edit from 'correcting' them to the usual 0.1x is a comment, and GPT-6.1 Sol does not have one. CE065 compares only the mirror with pricing.py, not the values themselves, so a 'correction' passes every gate. A small table-driven test in tests/test_pricing_registry.py that asserts the expected off-norm ratio for each listed id would catch this. (trigger: src/coder_eval/pricing.py) (restates: Axis 1: gpt-6.1-sol row has a 0.05x cache-read ratio with no annotation)

Nightly pipeline:

  • 🔵 The PR description does not say what happens to the nightly run. For nightly rows on these three models, cost changes from null to a number and cost_complete changes from false to true, which is a visible change in run.json and the board totals. Any nightly task that sets max_usd is now gated where before it was skipped, so its final_status can change to COST_BUDGET_EXCEEDED. Claude Code rows on Haiku 5.5 will also book about 40x cost until the CLI is bumped. The schema, the task.json shape and the CLI stay the same. The pricing change ships only with a lockstep image rebuild at release. Say this in the PR so the coder-eval-uipath nightly owners can expect the change in cost columns. (trigger: src/coder_eval/pricing.py) (restates: Axis 8: Haiku 5.5 is priced only at its <=100K-prompt tier)

Harness & Lint Improvements

Static checks (lint / type):

  • [ce-lint] CE069 (next free number; CE067 is unused but skip it to avoid confusion). Add a whole-tree @pytest.mark.lint class TestCE069OffNormCacheReadDeclared in tests/test_custom_lint.py, with its table OFF_NORM_CACHE_READ: dict[str, str] (key -> reason) beside DELIBERATELY_UNMIRRORED in tests/lint/pricing_mirror.py. Use builtin_rates() and leave out per_request_billing rows and DELEGATE_MODEL_IDS (that block has its own notes pointer). Every row whose cache_read_per_mtok / input_per_mtok != 0.1 must be a key of the table. The reason must contain the ratio literal (for example '0.05x'), so that it states a rule and not a history. Every key must still be priced and still off-norm, using the same stale-entry check as _assert_exemptions_are_live. Then delete the free-text ratio guard comments that the table replaces (pricing.py:63-64 Fable, :67 Opus 5.5, :105 codex-mini, :108 Pro tiers). A run on pr-220 gives these off-norm, non-exempt rows: claude-fable-5-1, claude-opus-5-5, claude-sonnet-5-5, claude-3-haiku-20240307 (0.12x, which has no annotation before this PR either), codex-mini-latest, gpt-5.5-pro, gpt-5.4-pro, gpt-6.1-sol, the three Bedrock rows and the two jev rows. Prevents: Finding 1: gpt-6.1-sol at pricing.py:126 has a 0.05x cache-read ratio with no annotation. The rule fails until the row is declared. Finding 2, in part: the Sonnet 5.5 row at pricing.py:81 must get a reason that contains '0.05x', which puts the ratio as a rule and not as a dated change. It also finds the old claude-3-haiku-20240307 0.12x row, which has no annotation.
  • [ce-lint] CE070: put tiered pricing in data, not in comments. Add tier_threshold_tokens: int | None = None (or flat_rate_floor: bool = False) to ModelPricing in src/coder_eval/pricing.py, and set it on every rate that has a long-context tier. Add a lint class in tests/test_custom_lint.py that fails on any own-line comment in src/coder_eval/pricing.py that matches #.*\b(tier|surcharge|long-context)\b|>\s*\d+K. After that, the only place a tier can be recorded is the field. Today the regex matches gpt-5.5 (:110), gpt-6-astra (:124), the Gemini Pro block (:127) and, on pr-220, claude-haiku-5-5 (:87). Delete those CAVEAT comments when you set the field. Also show the field in the evalboard mirror (render_pricing) so that the board can mark such a cost as a floor. Prevents: Finding 3: Haiku 5.5's 5x tier above 100K tokens is recorded only in a Python comment at pricing.py:87-88, so no code can act on it. It also removes the scoping problem in Finding 1: gpt-6.1-sol sits under gpt-6-astra's group CAVEAT comment, and no group comment can be read the wrong way when the tier is a per-row field.
  • [ci-gate] Extend make docs-budget (tests/lint/prose_budget.py) with a 'no dated history' rule. An own-line comment under src/ that has an ISO date (\b20\d{2}-\d{2}-\d{2}\b) must also have an expiry or contract marker (through|until|re-check|expires|HAZARD). There are no false positives on main today: the only dated comments, pricing.py:117 ('promotional through at least 2026-11-21; re-check then') and pricing.py:128 ('discounted by half through 2026-12-31'), are expiry markers. Known blind spot: history without a date ('used to', 'was') needs semantic judgment and stays a review item. Prevents: Finding 2: pricing.py:80 '5.5's cache hits fell to 0.05x input on 2026-10-07' is a dated change history, and CLAUDE.md says a comment states the contract, not the history. The new rule fails on this line.

Harness improvements (not statically reachable):

  • When a model has the tier field from CE070 and its cost comes from the rate card (_reprice_for_litellm claude_code_agent.py:1513, _backfill_cost :1494, evaluation/judge_usage.py:59, the simulator cost path), set cost_complete=false in run_record.py _cost_complete. Also emit the existing one-shot max_usd warning (orchestrator.py:1169-1178) that says the budget check is a lower bound. Add a unit test: put a turn with a prompt over 100K tokens on claude-haiku-5-5 through the LiteLLM reprice path, then check that cost_complete is False and that the warning is emitted. Why not static: Whether a prompt goes over the tier boundary is known only at runtime. calculate_cost also gets token buckets summed per turn, not per request, so no AST or type check can tell if a priced figure is a floor. Prevents: Finding 3: before this PR the model had no price, cost was None and cost_complete was false, which is honest. After this PR, a cost up to 5x too low is booked with cost_complete=true, and BudgetExceededError("usd") can fail to fire when the real bill is over the cap.
  • Generate a short list of tiered models (models whose max_usd value is a floor) from the CE070 field into docs/TASK_DEFINITION_GUIDE.md § Run Limits. Use the same generation and drift-check method as make plugin-reference / make pricing-mirror, so the list cannot fall out of sync with pricing.py. Why not static: Static checks can stop drift in the generated list. But the decision to tell users that max_usd is a lower bound is about wording and documentation, and only a review can confirm it. A static check cannot find a missing warning in prose until the list is generated. Prevents: Finding 3: today an author who sets max_usd on claude-haiku-5-5, gpt-5.5, gpt-6-astra or Gemini Pro is not told that the cap can read up to 5x low (2x for the GPT/Gemini tiers).

Top 5 Priority Actions

  1. Document in docs/TASK_DEFINITION_GUIDE.md under Run Limits that max_usd is only a lower bound for tiered models. Haiku 5.5 is the main case: at src/coder_eval/pricing.py:87-88 it is priced at its <=100K-prompt rate, but past 100K every bucket bills at 5x, so BudgetExceededError("usd") can fail to fire on the LiteLLM reprice path (claude_code_agent.py:1513), the backfill path (:1494) and the judge path (evaluation/judge_usage.py:59).
  2. Add a tiered (flat-rate floor) marker to ModelPricing and set it on claude-haiku-5-5, gpt-5.5 and gpt-6-astra, so that check_pricing_coverage warns and cost_complete reports false when one of these models runs under max_usd. Today these rows produce a confident number that can be too low, where before they produced an honest null.
  3. Add a guard comment above src/coder_eval/pricing.py:126 that marks the off-norm rate, for example # 6.1 Sol prices cache hits at 0.05x input, not 6 Sol's 0.1x.. Also say whether the >272K-tier caveat on line 124 applies to gpt-6.1-sol, or move the row out from under that caveat.
  4. Change the Sonnet 5.5 comment at src/coder_eval/pricing.py:80 from a dated history (fell to 0.05x ... on 2026-10-07) to a rule in the style of line 67. If the effective date matters for re-pricing older runs, use the existing HAZARD convention (lines 118-120) instead.
  5. After the pricing.py edits, run make pricing-mirror and make verify so that CE065 confirms evalboard/lib/pricing.generated.ts matches the three new rows. Also tell nightly consumers that runs on these models now report a numeric cost instead of cost=null.

Stats: 0 🔴 · 0 🟠 · 0 🟡 · 3 🔵 across 8 axes reviewed.

@uipreliga uipreliga left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fix what you agree with and 🚢

@bai-uipath
bai-uipath merged commit b3ae960 into main Oct 7, 2026
21 checks passed
@bai-uipath
bai-uipath deleted the bai/pricing-haiku-5-5 branch October 7, 2026 21:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants