Skip to content

feat(tui): show token throughput - #42112

Closed
opencode-agent[bot] wants to merge 1 commit into
v2from
tps-tally
Closed

opencode-agent[bot] wants to merge 1 commit into
v2from
tps-tally

Conversation

@opencode-agent

@opencode-agent opencode-agent Bot commented Aug 12, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • record the first visible text timestamp and provider generation completion timestamp for each V2 assistant step
  • calculate per-request output speed as (visible output tokens - 1) / (last token - first token)
  • show tok/s for completed stop and length responses in assistant footers and turn-token diagnostics
  • reject tool-call, error, buffered, and sub-250ms samples instead of presenting misleading throughput

This is the V2 implementation of the token-throughput requests in #5374 and #6096. The metric follows the standard TPOT/output-speed definition: TTFT and tool execution remain outside the generation window, hidden reasoning tokens are excluded, and non-streaming or too-short responses report no score. Live event projections and replayed message projections retain the same timing fields.

Checks

  • bun test test/cli/tui/data.test.tsx test/cli/tui/session-rows.test.ts (packages/tui)
  • bun typecheck (packages/tui)
  • bun test test/session-projector.test.ts test/session-runner-tool-events.test.ts (packages/core)
  • bun typecheck (packages/core)
  • bun typecheck (packages/schema)
  • bun typecheck (packages/client)
  • generated client output stability check
  • Prettier and git diff --check
  • replayed one identical Python API benchmark through the corrected UI: Grok 4.6 Fast 29.0 tok/s, GPT-5.6 Sol 28.1 tok/s, Claude Fable 5 N/A because its final response was buffered into a 66ms sample

Requested by: @R44VC0RP (vogel via Slack)

@Enough1122

Copy link
Copy Markdown

AI code review — automated review for reference; please use your judgment.

  1. packages/tui/src/routes/session/rows.ts:364 — const tokens = message.tokens.output - 1 subtracts an unexplained magic number from the output count — future readers cannot tell whether this compensates for the first token emitted before time.started is recorded or is an off-by-one bug waiting to be "fixed" — add a short comment naming exactly what the dropped token represents.

  2. packages/core/src/session/runner/publish-llm-event.ts:498 — const generated = yield* DateTime.now is a lexical declaration sitting directly in a case clause of the shared switch scope — it compiles, but it trips no-case-declarations style rules and leaks the binding to sibling cases — wrap the case "step-finish" body in braces or hoist the declaration above the switch.

  3. packages/tui/src/context/data.tsx:587 — the retry-reset (clearing time.started/time.generated/time.completed) and the started ??= event.created rule under session.text.started now duplicate the logic in packages/core/src/session/message-updater.ts:198/265 — two hand-synced copies can drift (one side clears a field the other forgets) and produce server-vs-client display mismatches — derive both from a shared helper or leave a cross-reference comment at each site.

  4. packages/tui/src/routes/session/index.tsx:1325 — the per-step usage table always renders the "Tok/s" column even when every row prints "-" because no step qualified — this wastes scarce terminal width in a UI that already gates content on dimensions().width — hide the header and cells whenever summary().throughput === undefined.

  5. packages/tui/src/routes/session/rows.ts:367 — the duration < 250 cutoff is a bare literal whose unit (ms) and rationale are implicit; misreading it as seconds would silently disable the metric — extract a named constant like MIN_THROUGHPUT_WINDOW_MS = 250 with a comment saying it suppresses noisy rates for very short generation windows.

  6. packages/core/src/session/runner/publish-llm-event.ts:498 — capturing generated before yield* flush() is deliberately right (it anchors the window at provider-finish instead of post-flush bookkeeping), but nothing records that intent — add a one-line comment so a later refactor does not reorder it and quietly skew every displayed rate.

Otherwise solid: the new optional schema fields keep old persisted messages parseable, the defensive if (!assistant) return guards fix latent crashes in the tool/reasoning/text-started handlers, and the helper is well tested (short-window, tool-call-step, and full event-pipeline cases).

@Enough1122

Copy link
Copy Markdown

AI code review - automated review for reference; please use your judgment.

  • packages/tui/src/routes/session/rows.ts:360-369 - tokens = output - 1 assumes the first token lands at time.started, but if providers include reasoning tokens in output (there is a separate reasoning field), the rate blends visible + hidden generation while tests call it "visible output throughput". Confirm the semantics or subtract reasoning too.
  • packages/core/src/session/runner/publish-llm-event.ts:498-505 - generated is captured at step-finish receipt time (before flush); under bus backpressure this slightly understates true completion time - acceptable for a display metric, worth a comment.
  • packages/core/src/session/message-updater.ts:198-199/254 - retry resets of started/generated plus the ??= on text.started keep the window idempotent across replays - good.
  • Weighted summary aggregation (sum tokens / sum durations) instead of mean-of-rates is the right call; guards against tiny windows (<250ms) avoid noisy values.

@github-actions

Copy link
Copy Markdown
Contributor

Automated PR Cleanup

Thank you for contributing to opencode.

Due to the high volume of PRs from users and AI agents, we periodically close older PRs using automated criteria so maintainers can focus review time on the most active and community-supported contributions.

This PR was closed because it matched the following cleanup criteria:

  • The PR was created more than 1 month ago
  • The PR had fewer than 2 positive reactions
  • Positive reactions are counted as thumbs-up, heart, celebration, or rocket reactions on the PR

PRs created within the last month are not affected by this cleanup.

If you believe this PR was closed incorrectly, or if you are still actively working on it, please leave a comment explaining why it should be reopened. A maintainer can review and reopen it if appropriate.

Thanks again for taking the time to contribute.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants