Problem
Every user turn wakes the full System Two LLM to make decisions that do not need a large generative model: is this a chat message or does it need a tool? can it be answered from context already in the conversation? how urgent is it? These are cheap classification decisions, but today they ride on the same expensive, slow, autoregressive call as the real work. On local models the cost is felt twice: latency and tokens spent on routing that a tiny model could answer in a single forward pass.
Proposed feature: a System One "Decision Gate"
A small, non-autoregressive System One decision model sits at the front door of the agent loop. It reads the pending turn and returns typed, calibrated answers (choice / noul / score) in one forward pass, without generating text. The harness can use those answers to skip expensive work on obvious requests.
This is the "Decision Gate" concept documented (CC BY 4.0, author Andrea Bruno) at harness-superfast. The idea is deliberately not patented and given to the community; attribution is the only condition.
Why this fits the existing architecture (architectural coherence)
The gate is designed to look and behave like features Qwen Code already ships, so it is a natural extension rather than a foreign object:
- Install-the-backend-via-a-command. Like
qwen sandbox inspects/proves its confinement backend, a new von-install subcommand provisions the decision backend out-of-band. The fork never bundles the model.
- A specialized decision model already exists here. The AUTO approval-mode classifier (
packages/core/src/permissions/classifier.ts) is a two-stage, schema-typed, fail-closed decision model. The Decision Gate is the same pattern applied at the front door.
- Off by default, behind a flag + a slash command.
/superfast on|off|status, persisted in settings.json under a new superfast section. Users who never run it see no change.
- Fails open. Any error, timeout, non-2xx, or low confidence returns "no opinion" and the harness behaves exactly as if the feature were off. The gate can only make things faster, never change behavior when unsure.
- No new runtime dependencies. The gate talks to a local
/v1/systemone endpoint over plain fetch.
Backend: Von
Von (wfzyx/von, Apache-2.0) is chosen because it is small (~1.5 GB) and runs on CPU, CUDA, ROCm, Apple Metal, and Intel iGPU — so the option is accessible on essentially any machine, including laptops with no discrete GPU. It is protocol-compatible with the Jev /v1/systemone spec, so any compatible server can be swapped in.
Scope of the first increment
The first PR ships the gate infrastructure running in shadow mode: when enabled, it classifies each turn and logs the recommendation, but does not yet change routing. This is the safe way to validate the model's accuracy against real traffic before acting on the decision. Acting on the route (skipping work) is a clearly-scoped follow-up once validated.
Alternatives considered
- Doing nothing (status quo): every turn pays full LLM cost for routing.
- A larger always-on gate: rejected as too invasive and unsafe without a validation phase.
Credits & License
The Decision Gate concept and architecture are by Andrea Bruno.
- Original white paper and license: https://github.com/Andrea-Bruno/harness-superfast
- Released under CC BY 4.0, intentionally not patented — given to the community; attribution is the only condition.
- The decision model (Von,
wfzyx/von) is third-party, Apache-2.0.
Please keep the attribution to Andrea Bruno and a link to the harness-superfast project in any reuse of this concept.
Problem
Every user turn wakes the full System Two LLM to make decisions that do not need a large generative model: is this a chat message or does it need a tool? can it be answered from context already in the conversation? how urgent is it? These are cheap classification decisions, but today they ride on the same expensive, slow, autoregressive call as the real work. On local models the cost is felt twice: latency and tokens spent on routing that a tiny model could answer in a single forward pass.
Proposed feature: a System One "Decision Gate"
A small, non-autoregressive System One decision model sits at the front door of the agent loop. It reads the pending turn and returns typed, calibrated answers (choice / noul / score) in one forward pass, without generating text. The harness can use those answers to skip expensive work on obvious requests.
This is the "Decision Gate" concept documented (CC BY 4.0, author Andrea Bruno) at
harness-superfast. The idea is deliberately not patented and given to the community; attribution is the only condition.Why this fits the existing architecture (architectural coherence)
The gate is designed to look and behave like features Qwen Code already ships, so it is a natural extension rather than a foreign object:
qwen sandboxinspects/proves its confinement backend, a newvon-installsubcommand provisions the decision backend out-of-band. The fork never bundles the model.packages/core/src/permissions/classifier.ts) is a two-stage, schema-typed, fail-closed decision model. The Decision Gate is the same pattern applied at the front door./superfast on|off|status, persisted insettings.jsonunder a newsuperfastsection. Users who never run it see no change./v1/systemoneendpoint over plainfetch.Backend: Von
Von (
wfzyx/von, Apache-2.0) is chosen because it is small (~1.5 GB) and runs on CPU, CUDA, ROCm, Apple Metal, and Intel iGPU — so the option is accessible on essentially any machine, including laptops with no discrete GPU. It is protocol-compatible with the Jev/v1/systemonespec, so any compatible server can be swapped in.Scope of the first increment
The first PR ships the gate infrastructure running in shadow mode: when enabled, it classifies each turn and logs the recommendation, but does not yet change routing. This is the safe way to validate the model's accuracy against real traffic before acting on the decision. Acting on the route (skipping work) is a clearly-scoped follow-up once validated.
Alternatives considered
Credits & License
The Decision Gate concept and architecture are by Andrea Bruno.
wfzyx/von) is third-party, Apache-2.0.Please keep the attribution to Andrea Bruno and a link to the
harness-superfastproject in any reuse of this concept.