Skip to content

Two-lane model strategy + plugin parity (port from claude-sdlc-wizard v2.0.0) #183

Description

@BaseInfinity

What claude-sdlc-wizard v2.0.0 just shipped

Two lanes with npm dist-tags

npm install agentic-sdlc-wizard              # Reliable (@latest) — tried and true
npm install agentic-sdlc-wizard@frontier     # Frontier (@frontier) — newest models

Plugin install (same result, no npm):

claude plugin install sdlc-wizard@sdlc-wizard-marketplace

The portable pattern (builder/reviewer/advisor + 95% escalation ladder)

Role What it does When
Builder Writes code, runs tests Always (it's you)
Reviewer (brain 1) Cross-model adversarial check When builder <95% confident, or before every merge
Advisor (brain 2) Deepest thinker, full context Only when builder + reviewer can't reach 95% together

Escalation ladder:

  1. Builder tries first. 95%+ confident → ship. No brain call.
  2. <95% → ask brain 1 with what you learned. Different model family = different blind spots.
  3. Still <95% or disagreement → ask brain 2 with what BOTH learned. Reconcile all three.

Each rung passes forward what it learned. No model redoes work.

claude-sdlc-wizard model table

Role Reliable (@latest) Frontier (@frontier)
Builder Opus 4.6[1m] max (Anthropic) Opus 5.5 xhigh (Anthropic)
Reviewer (brain 1) GPT-5.6 Sol xhigh (OpenAI) GPT-5.6 Sol xhigh (OpenAI)
Advisor (brain 2) Fable 5.1 high (Anthropic) Fable 5.1 high (Anthropic)

codex-sdlc-wizard — THE INVERSE

Builder is OpenAI. Reviewer is Anthropic (cross-family diversity). Families flip:

Role Reliable (@latest) Frontier (@frontier)
Builder GPT-5.6 Sol xhigh (OpenAI) ??? (OpenAI, newest)
Reviewer (brain 1) Opus 4.6 or Fable 5.1 via claude -p ??? (Anthropic)
Advisor (brain 2) ??? (strongest available, full context) ???

What needs porting

  1. Two-lane dist-tag infrastructure — release.yml with scripts/derive-dist-tag.sh (already proven in claude-sdlc-wizard #713). Prerelease versions get --tag frontier, stable gets --tag latest.

  2. Model table in all shipped docs — exact model IDs, exact effort levels, no aliases. "GPT-5.6 Sol xhigh" not "Sol" or "Codex high".

  3. Brain escalation ladder — documented in wizard doc, skills, tested in tests/test-escalation-ladder.sh (live E2E, isolated sessions). 95% confidence threshold documented in 3+ shipped files.

  4. Plugin parity (#725 on claude-sdlc-wizard) — codex --plugin should give the same experience as npm install.

  5. Model pin tests — tests/test-model-pins.sh (6 assertions: frontier model present, reviewer present and defaulted, advisor present, no stale refs).

Reliable = what you actually use

Same principle: Reliable is what the maintainer dogfoods. Frontier tracks newest models for experimenting. Owner will figure out the Codex-side model pins interactively.

Reference

  • claude-sdlc-wizard v2.0.0 (just shipped)
  • claude-sdlc-wizard #725 (plugin parity)
  • claude-sdlc-wizard #720 (CI → local tests)
  • claude-sdlc-wizard tests/test-escalation-ladder.sh, tests/test-model-pins.sh

Activity

  1. changed the title [-]Two-lane model strategy: Reliable + Frontier (port from claude-sdlc-wizard)[/-] [+]Two-lane model strategy + plugin parity (port from claude-sdlc-wizard v2.0.0)[/+] on Sep 27, 2026
  2. BaseInfinity commented on Sep 27, 2026

    @BaseInfinity
    OwnerAuthor

    Lessons to port from claude-sdlc-wizard (session 2026-09-27)

    1. Model IDs have no point releases — use exact IDs

    Wrong Correct
    claude-fable-5-1 claude-fable-5
    claude-opus-5-5 claude-opus-5

    No .1 or .5 variants exist for any current Claude model. This issue's body carries the old wrong IDs — fix before porting.

    2. Per-lane reviewer split (the portable pattern)

    The reviewer model differs by lane. The claude wizard now does:

    • Reliable: GPT-5.5 xhigh (field-proven)
    • Frontier: GPT-5.6 Sol xhigh (newer)

    For codex-sdlc-wizard (families inverted):

    • Reliable: Opus 4.6 or Fable 5 via claude -p (field-proven Anthropic reviewer)
    • Frontier: whatever newest Anthropic model fits

    The point: the reviewer is not one-size-fits-all across lanes.

    3. Exhaust the driver before escalating (THE key principle)

    95% confidence is the escalation threshold, but sub-95% does NOT mean escalate immediately. The driver must actively try to reach 95% first:

    • Web search, read more files, grep the codebase, check git history
    • Each rung is expensive and more effective when it knows what was tried
    • When escalating: pass forward what you researched and what specific question remains
    • "I checked X, Y, Z and I'm stuck on this" >> "I'm not sure, please review"

    This must be explicit in the shipped skill, not just implied by the 95% threshold.

    4. Regression tests must prove they catch regressions (mutation testing)

    Every shipped surface (skills, hooks, plugins, docs) needs a test. But a green test suite is worthless unless it goes RED on the wrong version. Proven today:

    • Introduced wrong model ID → test caught it (FAIL)
    • Swapped lane reviewer → test caught it (FAIL)
    • Broke cowork/skill parity → test caught it (FAIL)
    • Found one direction NOT guarded → added the test

    The codex wizard should have the same: for every regression test, a mutation that proves it fires.

    5. Local test runner runs EVERYTHING

    No skip gates, no env vars, no API-key checks. One command, all suites. The maintainer has the subs — the runner assumes it. ./tests/run-local.sh went from 6 suites to 84 (2,602 checks, 0 failures).

    6. GPT-6 models are live in Codex CLI

    Codex 0.157.1 shows GPT-6-Astra (default), GPT-6-Sol, GPT-6-Luna. GPT-5.5 is "Legacy." GPT-5.6-Sol is "Older." The codex wizard's lane tables need updating — model mappings TBD by maintainer with Codex directly.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions