A Claude Code plugin that helps cooperative Claude and Codex agents follow a reviewable workflow, avoid common mistakes, and understand why a step refused.
Generating code has become cheap. Verifying it, understanding it, and being able to say honestly what was checked has not. forge-plugin is built around that asymmetry: agent output is treated as a claim, and nearly everything durable in the system exists to test claims rather than to produce more of them.
It merges two systems — forge (gate chains, an adversarial review constitution, worktree discipline, a kill-switch, eval regressions) and codex-orchestrator (append-only execution logging and headless-agent tools) — into one arrangement for the Claude + Codex pair. The Claude main session orchestrates and verifies; fresh agents implement and review on routes resolved for each execution. The shipped defaults use Codex for implementation and first-pass review and Claude for the binding verdict. Support for other harnesses is deliberately out of scope.
Forge assumes that agents cooperate with the documented workflow. Its hooks, checks, and refusal messages make the intended sequence easier to follow and mistakes easier to diagnose; they are not a tamper-proof boundary against a process with the operator's OS authority.
Separate the author from the judge, and record the route. Implementers and
reviewers use distinct fresh sessions, and the binding reviewer remains distinct
from the author. A run freezes each role's resolved Claude-or-Codex route, and Forge
records the actual provider, model, and effort. Codex reviewers are OS-sandboxed
read-only; Claude reviewers are instruction-bounded and execution-capable. Cross-model
separation can reduce correlated mistakes when the routes differ, but it is recorded
evidence rather than guaranteed by construction; a same-model binding review is
surfaced as a historical-routing finding by tests/test_repo_conformance.py::check_run
(routing-design decision 12).
Trust nothing that was merely reported. Agent handoffs are claims, not evidence. Within a Forge workflow, the orchestrator re-runs each required gate in its own environment before authorizing a commit. A run's history records what was actually observed — command output, exit codes, SHAs read from git rather than remembered.
Fail closed. Gates are unconfigured until /forge:init fills them, and an
unconfigured gate refuses rather than passes. For recognized direct Git commands,
a Claude Code PreToolUse guard helps a cooperative agent avoid committing without
the expected review-backed authorization. It is a mistake-prevention backstop, not
a security boundary.
Exercise the mechanical controls. Load-bearing mechanical controls ship with tests designed to fail when the control is disabled. A test that still passes with its control removed is treated as a defect, not as coverage. Prompt-text checks are described separately because instruction presence is not model compliance.
- Claude Code ≥ 2.x
- Python ≥ 3.10 is the intended runtime target (standard library only — no packages, no build step)
- OpenAI Codex CLI on
PATH(codex --version), authenticated - bash on Linux or macOS (BSD userland) is the intended shell target.
Reintegration locking needs no
flockbinary: the worktree-merge skill holds the portable Git-common-dir arbiter through the Forge CLI (common-lock hold), which takes its optional kernel layer through Python'sfcntl. Only the commit-lock helper pair (acquire-commit-lock.shandrelease-commit-lock.sh, which also guard the decision-event lock) needsflockorlockf.
Repository CI currently proves Ubuntu with Python 3.13. macOS and the broader Python 3.10+ matrix are intended but currently unproven; portability fixes are tracked separately.
The repository is its own plugin marketplace. From any Claude Code session:
/plugin marketplace add nixlim/forge-plugin
/plugin install forge@forge
/reload-plugins
(For a local checkout, pass the absolute path to the repo instead of the GitHub slug.) Restart the session afterwards so the skills and hooks load. You should see seven skills:
| Skill | Purpose |
|---|---|
/forge:init |
Install or refresh the per-repo layer (region file, AGENTS.md splice, .codex/, evals) |
/forge:workflow |
Run a full orchestration lifecycle (plan → tasks → close → report) |
/forge:orchestrate |
One focused Codex-agent execution, review, or verification cycle |
/forge:report |
Author the final report from the run log and chain evidence |
/forge:commit |
The five-step fail-closed commit gate chain |
/forge:worktree-merge |
The four-gate merge chain with locked-rebase reintegration |
/forge:drift |
Mechanical drift sensing, then an operator-invoked periodic semantic review |
Installing also registers a PreToolUse mistake-prevention guard for recognized direct Git invocations, an advisory PostToolUse invariant guard, a Stop union that independently runs telemetry aggregation and the drift-staleness nudge, and a SessionStart nudge. Every Stop and SessionStart member is silent and inert outside a forge-initialized repository.
In the repository you want to govern:
/forge:init
Init confirms the project name and default branch; installs forge-project.md
(with its configuration regions rendered into both CLAUDE.md and AGENTS.md), the
.codex/ layer (agent routing plus an execpolicy deny-list), the gitignore block,
and .forge/ state directories; mines your CI configuration and git history to
propose gate commands; seeds and baselines the eval suite; then presents the whole
install diff for your explicit approval. It never commits on its own.
Until the regions are filled, gates fail closed — a merge stops with
forge: gate1-test-command not configured — run /forge:init.
If the repository already carries an older non-plugin forge installation, init detects it and migrates: regions are salvaged byte-identically from their original locations, eval fixtures are imported with their committed baselines rather than re-recorded (re-minting would launder an existing reviewer regression into a fresh "correct" result), and everything left behind is enumerated in a committed migration report.
Each layer answers a different question, and each is configured per repository rather than assumed.
| Layer | Question it answers |
|---|---|
| Gate 1 / Gate 2 | Do the project's own tests, linters and type checks pass? |
| Invariants | Do the repository's declared structural rules still hold? Every declared invariant must be an executable command; one that cannot be scripted is moved to a review bullet explicitly. |
| Assertion sensor | Does this test actually assert anything? Blocks confidently assertion-free Python tests; advisory where it cannot resolve the delegation. |
| Scoped mutation | Do the tests detect a deliberately broken implementation? Runs on changed files after Gate 1, advisory, with a documented path to becoming blocking once its cost and baseline are known. |
| Adversarial review | Does the change do what was intended, and what did the last reviewer miss? Eight baseline lenses plus a per-artefact profile; binary PASS/BLOCK, no hedging. |
| Risk tiers | How much scrutiny does this diff deserve? Derived from the diff at gate time against committed policy, promote-only, with a non-narrowable floor for control-class paths. |
| Drift sensing | Has anything decayed since the last change? Evals, gates on a clean tree, invariants, full mutation, category coverage, region staleness, telemetry. |
Both the commit chain and the merge chain end in a binding review, and the merge chain's final gate is mandatory even when every constituent commit took the fast tier.
- Kill-switch — create
AGENT_HALT(or scopedAGENT_HALT_commit) at the main checkout root. The Claude CodePreToolUseguard parses recognized directgitinvocations and selected environment-prefix forms while the sentinel is present. It inspects but never executes the submitted command;bash -cwrappers,command git, and Git alias configurations are outside its coverage by design. The guard assumes cooperative agents and prevents mistakes rather than providing tamper-proof enforcement. Agents are instructed never to create or clear sentinels. - Control-class changes — within the Forge workflow, gates, the constitution,
agent routing, hooks, evals and
forge-project.mdroute to the binding reviewer and wait for your explicit approval. - Audit — guard denials and halt detections append to
.forge/tmp/halt-audit.log; full orchestration history lives in the run journal under.codex-orchestrator/runs/<run-id>/. - Dead merge-lock owner — a reintegration holder killed outright leaves its
arbiter owner record (
agent-rebase.lockdirandagent-rebase.lock.intentin the Git common directory) and later entrants refuse until it is cleared. Forge instructions reserve clearing it for you after proving the recorded host and PID dead.
The append-only run journal records executions and decisions. The final report draws on that record and the chain evidence. Existing committed run archives remain readable history; Forge does not generate or require new ones.
Golden control regressions live in .forge/evals/tasks/. Gate-output and
verdict-derived fixtures take the exact prompt as their Input and cite chain ID,
candidate and evidence references as provenance; the committed .result is the
accepted verdict baseline. Routing,
reviewer, constitution or other control-class changes require a strict run:
STRICT=1 scripts/forge/run-evals.shAn empty suite exits non-zero. A gate that no fixture exercises is not satisfied by having no fixtures.
Point scheduled CI at the mechanical checker, with CLAUDE_PLUGIN_ROOT set to
the installed plugin root. It runs only the mechanical checker and does not invoke an
LLM. Scheduled CI never launches semantic review or any model; the semantic half
is operator-invoked:
on:
schedule:
- cron: "17 4 * * 1"
jobs:
forge-drift-mechanical:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
path: project
- uses: actions/checkout@v4
with:
repository: nixlim/forge-plugin
path: forge-plugin
- name: Run Forge mechanical drift checks
working-directory: project
env:
CLAUDE_PLUGIN_ROOT: ${{ github.workspace }}/forge-plugin
run: '"${CLAUDE_PLUGIN_ROOT}/scripts/forge/drift-check.sh"'When the Stop or SessionStart nudge reports a stale drift report, invoke
/forge:drift interactively for the semantic half. A CRITICAL drift finding
writes a block that refuses new runs until you clear it.
Several agents can work in one repository. Per-commit authorization is content-addressed by the staged diff, so two chains can hold authorization at once and, for commands the guard evaluates, one candidate's marker cannot admit another. Journals record an owner and refuse an append from a live foreign owner. Runs are admitted by declared file scope rather than refused outright, and overlapping scopes are named on refusal. Reintegrations serialize through one arbiter in the Git common directory, so worktrees using that common directory contend on the same lock.
Decision-event emission is designed for lossless concurrent appends on local POSIX filesystems. Current CI evidence covers Ubuntu; macOS is an intended but unproven target. NFS and SMB are unsupported, and Windows is out of scope.
python3 -m unittest discover -s testsThe plain command above is the sequential developer form. The committed gate-1
cell in forge-project.md discovers the same modules and runs them as a work queue
inside one cell: each module is its own unittest process, pulled longest-first by
max(1, min(8, os.cpu_count() or 1)) workers, after taking the host gate slot
(/dev/shm/agents-sem/gate/slot-0.lock) and waiting out CPU pressure for at most
300 seconds; both guards are no-ops where those paths do not exist. Measured on the
shared twelve-core host at load 3 to 5: 102 modules, 2,031 tests, 335 to 350 seconds
wall; it is not a promised speedup. A candidate whose every classified path is
docs-class records a gate-1 skip instead of running the cell; the docs-contract stack validation
still runs the prose-contract modules for it.
The Python entry point scripts/forge/cli.py is a compatibility shim over the
implementation modules under scripts/forge/forge_cli/.
The suite is stdlib only; the work-queue discovery finds 2,031 tests. Its
prose-contract tests are instruction-presence and textual-consistency checks: they
verify that required instructions exist, not that a model follows them. The release
end-to-end test checks plumbing with scripted verdicts, including a literal PASS;
it does not run a live reviewer. UPSTREAM records both vendored upstream SHAs and
every deliberate deviation.
- Design decisions: docs/design/0001-founding-decisions.md and docs/design/0002-verification-expansion.md
- Full specification: docs/specs/forge-plugin-spec.md
- Orchestration contract: docs/orchestration-contract.md
Upstreams (assessed cherry-pick, no automatic sync):
- https://github.com/nixlim/forge (itself vendored from mock-server/mockserver-monorepo, Apache-2.0)
- https://github.com/alexzh3/codex-orchestrator (via the nixlim fork)
MIT. See LICENSE.