Repository navigation
Stop classifying a passing test's title as a secrets failure (#379) - #380
Merged
Merged
Conversation
A TAP report prints one `ok N` line per assertion that passed, and that
line's description is whatever the test named itself -- application
vocabulary, not a runner diagnostic. discourse/discourse's `Plugins QUnit`
run (29572043439) failed on six `not ok` chat-component timeouts
(`# fail 6`), but PatchRail answered `secrets_or_permissions_failure` at
0.53 on its one witness: a test that GREEN-passed, whose scenario is an
"insufficient permissions" poll UI case:
ok 1523 [564 ms] - poll - Acceptance: Poll Builder - polls are
disabled: regular user - insufficient permissions
Read verbatim it sent a maintainer to `gh secret list` over the title of a
passing test. A signal matched only inside a passing TAP line now witnesses
nothing (added to the mention-only bounds), so the log lands on `unknown` --
the honest ceiling, since PatchRail has no browser-QUnit test class. A real
permissions failure is untouched: `insufficient permissions` on an error
line, `Resource not accessible by integration` and a denied github-actions
push all still carry the rule, because none arrives on a green `ok` line.
Zoo benchmark unchanged (223/223, top-1 1.0); the other 12 real-world logs
classify identically. Adds tests/test_tap_pass_line_not_secrets.py, the
committed run excerpt, a benchmark-doc row and a CHANGELOG entry.
PabloCodes7
force-pushed
the
fix/tap-pass-line-not-secrets
branch
from
July 17, 2026 10:41
28310bb to
e6957ff
Compare
PabloCodes7
added a commit
that referenced
this pull request
Jul 17, 2026
Extend the real-world benchmark from thirteen to sixteen public CI logs, filling three mainstream ecosystems it had no row for: mastodon/mastodon (an RSpec system spec timeout -> ruby_bundle_failure), phoenixframework/phoenix (mix test ExUnit failures -> elixir_mix_failure), and signalapp/Signal-Android (a Gradle validateDebugScreenshotTest task -> java_build_failure). All three classify correctly and identically across 0.6.1, 0.7.3 and main, so they force no code change. Committed as excerpts (Phoenix's caret-notation ANSI stripped for legibility) and measured with reproducible commands, matching the existing before/after rows. Also lists the discourse excerpt that #380 added. Benchmark 223/223 top-1 1.0, full suite green. Co-authored-by: PabloCodes7 <[email protected]>
Merged
PabloCodes7
added a commit
that referenced
this pull request
Jul 21, 2026
Cut 0.7.4 from main. The last release, 0.7.3 (2026-07-16), predates ten false-positive fixes that are sitting on main and not on PyPI: an installed pytest read as a failing test run (#376), a passing test's title read as a secrets failure (#380), a PHPUnit assertion read as a Composer failure (#378), Flutter's cached Gradle Wrapper (#382), Cabal's dependency resolution read as Maven (#383), Crystal and dune's make targets read as C/C++ (#384, #385), a parenthesized 504 (#391), mypy in a pixi manifest (#390), an unset TERM read as a missing secret (#392), a warning-only yarn install and a recovered checkout (#393), and a verdict held up by invocations alone reporting diagnosis-level confidence (#395). Every one of those is a wrong answer a maintainer gets today from pip install patchrail. Version bumped in the four places that spell it out (pyproject, __init__, README quickstart, uv.lock), CHANGELOG's Unreleased section dated, and the real-world benchmark's release-status paragraph corrected: it no longer claims six fixes are unreleased. Co-authored-by: PabloCodes7 <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #379.
On a real failed run of discourse/discourse (
Plugins QUnit, run 29572043439), six QUnit chat tests timed out (# fail 6), but PatchRail answeredsecrets_or_permissions_failureat 0.53. Its one witness was a test that passed:insufficient permissionsis the scenario that poll test asserts the UI handles — the description on anokline (TAP for "this passed"), not a credential the job lacked. Read verbatim it sent a maintainer togh secret listover the title of a green test.The fix
A TAP report prints one
ok Nline per assertion that passed; its description is application vocabulary, not a runner diagnostic. This adds a passing-TAP-line bound to the mention-only machinery, so a signal matched only inside anok N …line witnesses nothing. With no browser-QUnit class, the log then lands onunknown— the honest ceiling ruff, svelte and Symfony land on.The run's real failures start
not ok(skipped by the line-start match) and carry nothing this rule keys on. A genuine permissions failure is untouched:insufficient permissionson an error line,Resource not accessible by integration, and a denied github-actions push all still carry the rule.Evidence (reproducible)
patchrail==0.6.1(before)secrets_or_permissions_failure0.53patchrail==0.7.3(PyPI today)secrets_or_permissions_failure0.53main(this PR)unknown0.15Adds
tests/test_tap_pass_line_not_secrets.py, the committed run excerpt atexamples/real-world/discourse-29572043439-excerpt.log, adocs/real-world-benchmark.mdrow, and a CHANGELOG entry.