Repository navigation
Cap confidence when invocations alone carry the verdict - #395
Merged
Merged
Conversation
An invocation proves a tool RAN, never that it failed. When every signal behind the winning class is an invocation, the deferral above has already looked for a rule that witnessed a real error and found none -- so what comes back is the name of the tool that happened to be running when the job died. The count-based score sold that as a diagnosis: rails/rails run 29648807728 dies on a Ruby SyntaxError that aborts bin/rails, no class covers it, and three Bundler invocations (bundle install -- which succeeded -- bundle exec, bundler) produced ruby_bundle_failure at 0.89, pointing a maintainer at a Gemfile that is fine. Such a verdict now caps at 0.3: still printed, because one lead beats none, but no longer sold as an answer. The guard is on the mechanism, not on any ecosystem or log -- a rule that actually watched something fail matches a signal outside INVOCATION_ONLY_PATTERNS and never reaches it, so legitimate last resorts (a bare pytest invocation) keep class and confidence both. No new class, no new zoo fixture: all 223 examples/ci-triage cases classify identically, class AND confidence, top-1 still 1.0. The regression guard is the raw rails log under tests/data/realworld/, pinning the confidence band.
Merged
PabloCodes7
added a commit
that referenced
this pull request
Jul 21, 2026
Cut 0.7.4 from main. The last release, 0.7.3 (2026-07-16), predates ten false-positive fixes that are sitting on main and not on PyPI: an installed pytest read as a failing test run (#376), a passing test's title read as a secrets failure (#380), a PHPUnit assertion read as a Composer failure (#378), Flutter's cached Gradle Wrapper (#382), Cabal's dependency resolution read as Maven (#383), Crystal and dune's make targets read as C/C++ (#384, #385), a parenthesized 504 (#391), mypy in a pixi manifest (#390), an unset TERM read as a missing secret (#392), a warning-only yarn install and a recovered checkout (#393), and a verdict held up by invocations alone reporting diagnosis-level confidence (#395). Every one of those is a wrong answer a maintainer gets today from pip install patchrail. Version bumped in the four places that spell it out (pyproject, __init__, README quickstart, uv.lock), CHANGELOG's Unreleased section dated, and the real-world benchmark's release-status paragraph corrected: it no longer claims six fixes are unreleased. Co-authored-by: PabloCodes7 <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
An invocation proves a tool ran. It never proves the tool failed.
When every signal behind the winning class is an invocation, the classifier has already deferred once — it looked for a rule that witnessed a real error and found none. What it hands back is the name of the tool that happened to be running when the job died: a lead, not a diagnosis. The count-based score sold it as one anyway.
Real case, rails/rails run
29648807728: the job dies on a RubySyntaxErrorinactionpackthat abortsbin/rails. No failure class covers that. The only matches in 3k lines are three Bundler invocations —bundle install(which succeeded),bundle exec,bundler— and three matches scoredruby_bundle_failureat 0.89, pointing a maintainer at a Gemfile that is fine.Fix
A verdict whose signals are all in
INVOCATION_ONLY_PATTERNScaps at 0.3. Still printed — one lead beats none — but no longer sold as an answer, and plainly above the0.15of a decline.The guard is on the mechanism, not on Ruby, Bundler, or this log: nothing in it names an ecosystem. A rule that actually watched something fail matches a signal outside the invocation set and never reaches the cap, so legitimate last resorts keep both class and confidence — a bare
pytestinvocation still answerspython_test_failureat 0.53.No new failure class, no new zoo fixture.
Verification
ci benchmark examples/ci-triage→ 223/223, top-1 = 1.0, unchanged.ruby_bundle_failure0.89 →ruby_bundle_failure0.3, same class, honest number.tests/data/realworld/(not the zoo — the benchmark count stays 223), pinning the confidence band rather than the class.