Skip to content

Stop classifying a passing test's title as a secrets failure (#379) - #380

Merged
PabloCodes7 merged 1 commit into
mainfrom
fix/tap-pass-line-not-secrets
Jul 17, 2026
Merged

PabloCodes7 merged 1 commit into
mainfrom
fix/tap-pass-line-not-secrets

Conversation

@PabloCodes7

Copy link
Copy Markdown
Contributor

Closes #379.

On a real failed run of discourse/discourse (Plugins QUnit, run 29572043439), six QUnit chat tests timed out (# fail 6), but PatchRail answered secrets_or_permissions_failure at 0.53. Its one witness was a test that passed:

ok 1523 [564 ms] - poll - Acceptance: Poll Builder - polls are disabled: regular user - insufficient permissions

insufficient permissions is the scenario that poll test asserts the UI handles — the description on an ok line (TAP for "this passed"), not a credential the job lacked. Read verbatim it sent a maintainer to gh secret list over the title of a green test.

The fix

A TAP report prints one ok N line per assertion that passed; its description is application vocabulary, not a runner diagnostic. This adds a passing-TAP-line bound to the mention-only machinery, so a signal matched only inside an ok N … line witnesses nothing. With no browser-QUnit class, the log then lands on unknown — the honest ceiling ruff, svelte and Symfony land on.

The run's real failures start not ok (skipped by the line-start match) and carry nothing this rule keys on. A genuine permissions failure is untouched: insufficient permissions on an error line, Resource not accessible by integration, and a denied github-actions push all still carry the rule.

Evidence (reproducible)

version verdict on the discourse excerpt
patchrail==0.6.1 (before) secrets_or_permissions_failure 0.53
patchrail==0.7.3 (PyPI today) secrets_or_permissions_failure 0.53
main (this PR) unknown 0.15
  • Zoo benchmark: 223/223, top-1 1.0 — unchanged.
  • The other 12 real-world logs classify identically before/after.
  • Full suite: 963 passed, ruff clean (Python 3.12).

Adds tests/test_tap_pass_line_not_secrets.py, the committed run excerpt at examples/real-world/discourse-29572043439-excerpt.log, a docs/real-world-benchmark.md row, and a CHANGELOG entry.

A TAP report prints one `ok N` line per assertion that passed, and that
line's description is whatever the test named itself -- application
vocabulary, not a runner diagnostic. discourse/discourse's `Plugins QUnit`
run (29572043439) failed on six `not ok` chat-component timeouts
(`# fail  6`), but PatchRail answered `secrets_or_permissions_failure` at
0.53 on its one witness: a test that GREEN-passed, whose scenario is an
"insufficient permissions" poll UI case:

    ok 1523 [564 ms] - poll - Acceptance: Poll Builder - polls are
      disabled: regular user - insufficient permissions

Read verbatim it sent a maintainer to `gh secret list` over the title of a
passing test. A signal matched only inside a passing TAP line now witnesses
nothing (added to the mention-only bounds), so the log lands on `unknown` --
the honest ceiling, since PatchRail has no browser-QUnit test class. A real
permissions failure is untouched: `insufficient permissions` on an error
line, `Resource not accessible by integration` and a denied github-actions
push all still carry the rule, because none arrives on a green `ok` line.

Zoo benchmark unchanged (223/223, top-1 1.0); the other 12 real-world logs
classify identically. Adds tests/test_tap_pass_line_not_secrets.py, the
committed run excerpt, a benchmark-doc row and a CHANGELOG entry.
@PabloCodes7
PabloCodes7 force-pushed the fix/tap-pass-line-not-secrets branch from 28310bb to e6957ff Compare July 17, 2026 10:41
@PabloCodes7
PabloCodes7 merged commit d320d15 into main Jul 17, 2026
5 checks passed
@PabloCodes7
PabloCodes7 deleted the fix/tap-pass-line-not-secrets branch July 17, 2026 10:47
PabloCodes7 added a commit that referenced this pull request Jul 17, 2026
Extend the real-world benchmark from thirteen to sixteen public CI logs,
filling three mainstream ecosystems it had no row for: mastodon/mastodon
(an RSpec system spec timeout -> ruby_bundle_failure), phoenixframework/phoenix
(mix test ExUnit failures -> elixir_mix_failure), and signalapp/Signal-Android
(a Gradle validateDebugScreenshotTest task -> java_build_failure).

All three classify correctly and identically across 0.6.1, 0.7.3 and main, so
they force no code change. Committed as excerpts (Phoenix's caret-notation ANSI
stripped for legibility) and measured with reproducible commands, matching the
existing before/after rows. Also lists the discourse excerpt that #380 added.

Benchmark 223/223 top-1 1.0, full suite green.

Co-authored-by: PabloCodes7 <[email protected]>
@PabloCodes7 PabloCodes7 mentioned this pull request Jul 21, 2026
PabloCodes7 added a commit that referenced this pull request Jul 21, 2026
Cut 0.7.4 from main. The last release, 0.7.3 (2026-07-16), predates ten
false-positive fixes that are sitting on main and not on PyPI: an installed
pytest read as a failing test run (#376), a passing test's title read as a
secrets failure (#380), a PHPUnit assertion read as a Composer failure (#378),
Flutter's cached Gradle Wrapper (#382), Cabal's dependency resolution read as
Maven (#383), Crystal and dune's make targets read as C/C++ (#384, #385), a
parenthesized 504 (#391), mypy in a pixi manifest (#390), an unset TERM read as
a missing secret (#392), a warning-only yarn install and a recovered checkout
(#393), and a verdict held up by invocations alone reporting diagnosis-level
confidence (#395). Every one of those is a wrong answer a maintainer gets today
from pip install patchrail.

Version bumped in the four places that spell it out (pyproject, __init__,
README quickstart, uv.lock), CHANGELOG's Unreleased section dated, and the
real-world benchmark's release-status paragraph corrected: it no longer claims
six fixes are unreleased.

Co-authored-by: PabloCodes7 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

A phrase in a passing test's title is classified as a secrets/permissions failure

1 participant