Repository navigation
Releases: patchrail/patchrail
Release list
v0.7.6
When PatchRail cannot name the cause, it now shows you where the log ends.
1,164 lines of apache/kafka run 29805964231 used to go in and this came out: unknown, 0.15, "No high-confidence local signal found." — not one line of the log. The cause was sitting one line above the runner's boilerplate exit-code annotation (Could not find the PR that triggered this workflow request), three seconds of scrolling away. Declining to guess is the right call; answering with nothing is not.
Added
- An
unknownverdict now hands back the last lines of the log that carried output, under a heading that says exactly what they are: raw log, not a diagnosis. The heuristic is positional, never lexical — it asks where the log stopped, a question with the same answer in every ecosystem, including the ones no rule covers. Optionallog_tailkey inci-result.v1(max 5 lines, each capped at 300 characters and passed throughredact_ci_log). (#399) - The same tail follows a verdict that is only a hint (under
LOW_CONFIDENCE_THRESHOLD= 0.35). rails/rails run 29648807728 is the case: threebundleinvocations carryruby_bundle_failureat 0.3, none of them failed, and the step that actually broke —./bin/rails assets:precompile— sat in the last lines nobody was shown. Same extraction, same cap, same redaction; only the condition widened. (#400)
failure_class, confidence, signals and exit codes are untouched, and a log that classifies produces byte-identical output to 0.7.5.
Full changelog: https://github.com/patchrail/patchrail/blob/main/CHANGELOG.md
v0.7.5
A low-confidence verdict now reads like one.
0.7.4 stopped a class carried by nothing but tool invocations from claiming diagnosis-level confidence, but it printed that capped verdict in exactly the same shape as a confident one — Root cause: <class> / Confidence: 0.3 — so a maintainer still read "root cause" and went off to fix a file that was fine. The number dropped; the message didn't.
Fixed
- Both human-facing reports (
--format textand--format markdown) now add an explicit caveat below a configurable threshold (LOW_CONFIDENCE_THRESHOLD = 0.35): that the verdict is a hint rather than a proven cause, that the signals behind it only show a tool ran and not that it failed, and that the next useful step is the raw log — plus the existing CI failure fixture issue template if the real cause turns out to be something PatchRail cannot classify yet. The guard keys on confidence alone, never on a class or ecosystem;unknownkeeps its own existing message instead of being told twice. (#397)
Presentation only: the class, the confidence, the exit code and the --format json payload the GitHub Action consumes are unchanged, and a high-confidence report is byte-identical to 0.7.4.
Full changelog: https://github.com/patchrail/patchrail/blob/main/CHANGELOG.md
v0.7.4
Ten false positives removed. Each one is a verdict 0.7.3 gets wrong on a real CI log — the classifier named a tool that was fine, or reported diagnosis-level confidence for a guess.
Fixed
- A verdict carried by invocations alone now caps at 0.3 confidence. An invocation proves a tool ran, never that it failed; rails/rails run 29648807728 dies on a Ruby
SyntaxErrorand threebundleinvocations were producingruby_bundle_failureat 0.89, sending a maintainer to debug a Gemfile that is fine. Legitimate last resorts (a barepytest) keep both class and confidence. (#395) - A yarn install that finishes with warnings is no longer a broken dependency install, and a
git checkoutthat recovers is no longer a checkout failure. (#393) debconf: (TERM is not set, ...)and other unset terminal/locale variables no longer read as a missing repository secret. (#392)- A parenthesized
Gateway Timeout (504)isnetwork_transient_failure, notunknown. (#391) - mypy listed in a pixi/conda dependency manifest is no longer a type-check failure. (#390)
- dune's
make testand a Crystal spec run are no longer read as C/C++ builds. (#384, #385) - Cabal's
Could not resolve dependencies:is no longer a Maven build failure. (#383) - Flutter's cached Gradle Wrapper is no longer a Gradle build failure. (#382)
- A passing test's title is no longer a secrets failure. (#380)
- A failing PHPUnit assertion is no longer a Composer failure. (#378)
- An installed
pytestis no longer reported as a failing test run. (#376) --outwrites only to the file, not also to stdout;redactandpilot-packerror on an empty log. (#370, #371)
Changed
patchrail --helpnow ends with a link to the project home.- The CI triage job summary shows the local reproduction command, and no longer carries internal adoption telemetry. (#372, #373)
- PyPI
keywordsand trove classifiers broadened to match how a CI-triage tool is actually searched for.
Two real mainstream CI logs (spring-boot, kafka) are now pinned as regression tests under tests/data/realworld/, and docs/real-world-benchmark.md covers twenty-one committed public logs.
Full changelog: https://github.com/patchrail/patchrail/blob/main/CHANGELOG.md
v0.7.3
Four new-user UX fixes merged since 0.7.2, now on PyPI.
Fixed
- Piping a non-UTF-8 CI log no longer crashes. The headline one-liner
gh run view <run-id> --log-failed | patchrail ci explainfeeds raw log bytes to stdin, and real CI logs often carry bytes that aren't clean UTF-8 (a latin-1 accent in a test name, stray ANSI/control bytes, a truncated multibyte sequence).explain,classify,pilot-packandredactnow read the stdin pipe as bytes and decode witherrors="replace"— the same forgiveness the--log <file>path already had — instead of raising an uncaughtUnicodeDecodeErroron a first-timer (#368). ci explainwith no--login a terminal no longer hangs. Runningpatchrail ci explainwithout piping anything in used to block forever onstdin.read(), looking frozen. It now detects an interactive terminal and fails fast with a hint to pass--logor pipe a log in, exit 2. A real pipe or file redirect is untouched (#366).--logat a directory or an unreadable file gives a clear message, not a traceback.patchrail ci explain --log logs/used to leak a rawIsADirectoryError/PermissionError. It now reports the problem on stderr and exits 2, the same clean contract as the missing-file case (#365).- A passing CI log is no longer reported as an unrecognized failure to file a fixture for. A green log that matched no failure rule landed on
unknown(0.15) with "Open a CI failure fixture issue" — inviting non-failures into the tracker. PatchRail now recognizes a log that plainly announces success, flagslikely_successful_run, and replies "No failure detected — point me at the failed run." Conservative: any failure tell vetoes it, so a genuinely unrecognized failure keeps itsunknownverdict (#364).
Also includes a GitLab CI clang++ undefined-reference link fixture from @sahilmathur254 (#363).
v0.7.2
Three fixes merged since 0.7.1, now on PyPI.
Fixed
- Flyway runtime migration failures are recognised. A
flyway migratestep that fails against the database —ERROR: Migration V2__add_users.sql failed,SQL State : 42S01— used to fall through tounknown(0.15). It now classifies asdatabase_migration_failureat 0.71 and points you at the SQL error to fix; theSQL State : 00000success code is explicitly excluded so green runs are untouched. Thanks to @hkJerryLeung for the fix (#361, closing #265) — PatchRail's first merged external contribution. - A failed
yarn runscript is no longer reported as a dependency install. yarn classic's docs footer was matched on the bare host, so ayarn run prettierthat failed on a diff got sent to reconcile a lockfile. The footer is now pinned toinstall/add; a realyarn installfailure is unaffected (#360). - A container OOM-killed under test is no longer reported as the runner running out of memory.
OOMKilled/ exit 137 /Out of memoryfrom a container runtime's own tests now defer to the concrete cause the log recorded (herego_test_failure). A real host exhaustion still trips a terminal signal and keeps its verdict (#358).
pip install --upgrade patchrail · patchrail --version → patchrail 0.7.2
v0.7.1
Two misreadings, both the same shape: a line that merely names a tool, read as proof that tool had failed.
Fixed
-
A job that declares
timeout-minutesis no longer reported as a job that timed out. Actions echoes your step config into the log —timeout-minutes: 180— on green runs and red ones alike, and PatchRail was reading that declaration as evidence the limit had been hit. On envoyproxy/envoy's coverage run29363920524the job never came near its 180-minute ceiling: one directory out of 430 had slipped three tenths of a point under its coverage threshold. That log now answerscode_coverage_threshold. A job that really did run long is unaffected — the runner says so in words (has exceeded the maximum execution time of…,The operation was canceled), and those still land asci_job_timeout. -
A Go job that fails to compile is no longer reported as a lint failure.
golangci-lint-actionnames its tool five times just to ship it — version lookup, cache hit, install, echoed command line, timing — and PatchRail was reading those install lines as proof the linter had failed. Any Go job that merely declared the action got a confidentgo_lintverdict, whatever had actually broken. On grafana/grafana'slint-gorun27635190952there was no lint finding at all: a function had grown a parameter and eight call sites in one test file had not. That log now answersgo_test_failureand reproduces withgo test ./.... A linter that really did report something ((gofmt),(gci),(revive)) is stillgo_lint. -
Go call-site mismatches are recognised.
not enough arguments in call toandtoo many arguments in call to— the compiler's own words — no longer pass unread, alongsideundefined:.
Both fixes were measured against the real failing logs, which ship in the repo under examples/real-world/ so you can reproduce the before and after yourself.
pip install --upgrade patchrail
patchrail ci explain --log examples/real-world/envoy-29363920524-excerpt.logFull changelog: https://github.com/patchrail/patchrail/blob/main/CHANGELOG.md
v0.7.0 — nine fixes for one mistake
Nine misreadings, all of one kind.
Every fix in this release started as a real failing run at a real project — pandas, deno, svelte, istio, astro, ruff, prometheus, pytorch — where PatchRail answered confidently and answered wrong. And every one of them was wrong in the same way: it mistook a line that merely named a tool for a line proving that tool had failed.
An install listing. A script echoed before it runs. A suppressible CMake warning. A cache cleaning up after a job that was already dead. A test quoting a compiler back at itself.
A verdict that no signal ever witnessed failing now yields to what the runner actually flagged — and when nothing is left, the honest answer is unknown, with the failing line handed back to you.
What stops being wrong
- A CMake policy that "is not set" is no longer read as a repository secret that is not set (
CMP0148is shaped exactly like an environment variable). pytorch's lint job died injq; we sent maintainers to audit their secrets. - A dependency Dependabot merely listed in its job definition is no longer mistaken for a linter that ran and failed. svelte's maintainers were told
pnpm lintbroke. No linter ran. - Output a test quotes back at you is no longer read as the job's own diagnostic. deno's spec suite typechecks deliberately-broken programs on purpose; we told the maintainers of a TypeScript runtime that their TypeScript was broken, at 0.95 confidence.
- The source of a step, echoed before it runs, is no longer read as the step's output. ruff's Rust panic came out
python_test_failure, on an error branch for a download that had succeeded. - A cache that could not save, and a tool the job merely installed, no longer decide the verdict. pandas' doc build sent a maintainer to debug a cache that was working perfectly.
- A failed
npm auditis reported as a failed scan, not a broken install — and npm's post-install audit tally is not a failed scan at all. - A tie between two failure classes goes to the rule that watched something fail. prometheus — a Go repo, whose Go tests failed — came out
javascript_lint. - A proxy logging its own client disconnects is not a network outage.
New: the classifier, measured against real logs
docs/real-world-benchmark.md runs PatchRail against seven real CI logs and shows what it said before, what it says now, and the command that reproduces each one. The logs are committed verbatim under examples/real-world/, so you can check the numbers instead of trusting them.
The misses are in the table, not in a footnote. pandas still lands on python_test_failure at 0.53 rather than naming the crash. grafana calls a compile error go_lint. Four of the seven verdicts are identical before and after — which is the point: these fixes are narrow, and a benchmark that only shows wins is not a benchmark.
pipx install patchrail
patchrail ci explain --log failed-ci.logRuns locally. No API keys, no log upload.
v0.6.1
Patch release: a success announced through the error channel is no longer reported as somewhere to "start".
oven-sh/bun's failing run carries exactly one runner annotation in 4,709 lines, and the workflow emits a success through it — ##[error]✅ Autofix task started. The unknown verdict handed that line back under "the CI runner did annotate these lines as errors — start there", pointing the maintainer at a line saying everything went fine.
The verdict itself was, and stays, correct: unknown is honest for that log. Only the evidence changes. An annotation is now dropped when it opens with a success mark and names no failure anywhere in the line — a guard deliberately lopsided towards keeping, so ✅ 2 passed, ❌ 1 failed and ✔ image built, but the upload failed both survive.
Users of the CI Triage Action pick this up on their next run with no change on their side.
Closes #329. Full notes in CHANGELOG.md.
v0.6.0 — an unknown log that explains itself, and no more invented diagnoses
Three changes that landed after 0.5.0, all found the same way: piping real failing runs from public repositories through the one-liner in the README, the way a first-time user meets this tool. Eight new ecosystems (deno, bun, istio, astro, airflow, envoy, ruff, pydantic) — three came back confidently wrong.
Fixed
- A tool named inside a filename is no longer a diagnosis. A CI log in a monorepo is mostly filenames, and a rule could be carried start to finish by a word that only ever appeared inside one.
oven-sh/bun— a Zig/JS runtime — was told its pnpm lockfile was out of date, on the evidence of prettier listing files it left(unchanged), two of which sat under alockfile/directory; in the same log, a regression test named after an advisory (GHSA-pfwx-36v6-832x.test.ts) was read as a failed security scan that never ran.istio/istio, a Go repository, was diagnosed with both a Java build failure and a Node install failure, read out of a 2929-character single-line Dependabot JSON blob.withastro/astromatched on--no-frozen-lockfile— the flag that permits the change it was complaining about. A signal must now match outside a path token to carry a verdict, and a bare tool name corroborates a failure without ever constituting one. - The runner's content-free annotations no longer dress an empty answer up as a finding.
Process completed with exit code 1.appears in every failing run, whatever the cause; reported back under a heading it read as evidence, and — each exit code being a distinct string — a matrix build's worth of them would crowd out the annotation that actually names the failure.
Added
- An
unknownverdict now hands back the line the runner itself flagged. A log no rule matches usually still names its own failure:psf/requestsfailed on a single, self-explanatory line —##[error]"github-token" length must be less than or equal to 100 characters long— and PatchRail answeredunknown, no signals, nothing else. Those annotations are now reported inrunner_errors(JSON), under "Errors the runner reported" (Markdown), and asRunner reported:(text), redacted and de-duplicated. They stay evidence and nothing more: an annotation says where the job died, not why, so the class staysunknownand the confidence stays at 0.15. PatchRail still does not pretend to recognize a log it does not recognize.
Genuine failures are unaffected: real lockfile, security-scan and build failures score exactly as before, and no fixture changes classification or confidence (221 fixtures, top-1 1.0).
pip install --upgrade patchrailv0.5.0 — real logs, real causes
Five accuracy fixes that landed after 0.4.0, all found by running ci explain over real failing runs from public repositories rather than over the fixture zoo. The zoo's logs are clean; a real one is not, and every one of these bugs hid in the difference.
Fixed
- A tool the job merely named no longer outranks the one that failed.
set -xechoes every command into the log whether it passes or not, and pip announces every dependency it resolves — soCollecting mypy==1.17.1put mypy in the log of a job that never type-checked anything. Signals are now judged by the line they land on. Caught 3 of 9 real Python runs (httpx, prefect, home-assistant) and 3 of 17 across Go/Rust/JVM/.NET/C++ (kafka, moby, opencv), where aGRADLE_HOME=entry in a Windows environment table was enough to diagnose a Java build failure in a job that compiled nothing. - ANSI colour codes no longer hide the failure. CI keeps the colour on and the GitHub log API serves the escapes back, so the reset lands inside the failure line (
^[[31mFAILED^[[0m tests/…) andFAILED .*::silently stopped matching. A real airflow test failure was reported as an artifact error. python_test_failurenow recognises pytest's own verdict —1 failed, 1416 passed in 18.37s, the short test summary, and collection errors, not just named tests.- Post-failure cleanup noise no longer outranks the real cause. The
if: failure()step that uploads logs and finds nothing is the commonest shape in CI; it used to tie the real cause and win on declaration order. - The classifier no longer reads a command that merely ran as the failure —
\btsc\bmatched the x86 time stamp counter in/proc/cpuinfo, diagnosing a rust-lang/rust build failure as a TypeScript typecheck.
Genuine failures are unaffected: a real gradle, clippy, docker or dotnet failure scores exactly as before, and no fixture changes classification or confidence (221 fixtures, top-1 1.0).
Added
patchrail schema ci-classespublishes the schema forci classes --format json, with a conformance test that pins the emittedschema_versionto the schema you can fetch, so the contract cannot drift silently again.docs/json-cookbook.mddocuments theci classespayload with a coverage recipe.
pip install --upgrade patchrail