Skip to content

chore(bench): shadow hook that sizes RTK's real output on read-only commands - #506

Merged
inth3shadows merged 1 commit into
mainfrom
bench-rtk-shadow
Oct 7, 2026
Merged

inth3shadows merged 1 commit into
mainfrom
bench-rtk-shadow

Conversation

@inth3shadows

Copy link
Copy Markdown
Owner

Why

#505 priced what RTK would reach of Bash output but could only give a ceiling (7.73% of the bill): rtk pipe over recorded output is not what the live rtk <cmd> prints. This measures the live command, without changing what any session sees.

What

scripts/bench/rtk_shadow.py, two entry points:

  • hook <dir>: a Claude Code PostToolUse hook for Bash. For a single simple read-only command (grep, rg, cat, head, tail, ls, find, git log/diff/status/show; a leading cd <absolute dir> && is allowed) it runs RTK's form once, detached, and appends a row with the token size of RTK's output beside what the session was shown. Every other command gets a row with a skip reason.
  • report <dir>: per-day summary. The saving is counted two ways: exit-0 RTK runs only (headline), and any exit.

Safety

  • Nothing off the read-only list is ever re-run; options that write or execute (find -delete/-exec, rg --pre, git --output, --help) are refused.
  • The RTK run has no network, its own pid namespace, a 4 GiB memory limit, a 15 s timeout, and a throwaway home on tmpfs (RTK keeps command history under HOME).
  • Rows hold sizes and a label from a fixed list of program names: no command line, path or output.
  • Hook mode always exits 0 and prints nothing.

Known limits (stated in the report footer)

  • Coverage is what the harness re-runs, which is less than what RTK would take.
  • Claude Code sends no hook event for a failed Bash call, so failed commands are in no row; the last two report columns are not a strict floor.
  • It prices what RTK removes, not whether the session then needed it.

Tests

uv run pytest scripts/bench/tests/ -q: 101 passed, 1 skipped. ruff check scripts/bench/ clean. Reviewed four times before and after install; the first pass blocked it (a tree -ao write bypass and a hang on an escaped child, both fixed and pinned by tests).

No change under src/; no release.

…ommands

The shell replay (#505) could only give a ceiling for RTK: `rtk pipe` over
recorded output is not what the live `rtk <cmd>` prints. This is the live
measurement, without changing what any session sees.

rtk_shadow.py is a Claude Code PostToolUse hook for Bash. For a command that
is one simple read-only command (grep, rg, cat, head, tail, ls, find, git
log/diff/status/show) it runs RTK's form once, detached, and logs the token
size of RTK's output beside what the session was shown. Every other command
gets a row saying why it was skipped, so coverage is reported. Rows hold
sizes and a label from a fixed list; no command line, path or output.

The RTK run has no network, its own pid namespace, a memory limit and a
throwaway home on tmpfs. The hook always exits 0. `report` prints a per-day
summary and counts the saving two ways (exit-0 runs only, and any exit).
@inth3shadows
inth3shadows merged commit f65f870 into main Oct 7, 2026
7 checks passed
@inth3shadows
inth3shadows deleted the bench-rtk-shadow branch October 7, 2026 21:28
@inth3shadows
inth3shadows restored the bench-rtk-shadow branch October 7, 2026 23:47
@inth3shadows
inth3shadows deleted the bench-rtk-shadow branch October 8, 2026 01:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant