Skip to content

Releases: patchrail/patchrail

v0.7.6

Choose a tag to compare

@PabloCodes7 PabloCodes7 released this 22 Jul 19:28
2c57412

When PatchRail cannot name the cause, it now shows you where the log ends.

1,164 lines of apache/kafka run 29805964231 used to go in and this came out: unknown, 0.15, "No high-confidence local signal found." — not one line of the log. The cause was sitting one line above the runner's boilerplate exit-code annotation (Could not find the PR that triggered this workflow request), three seconds of scrolling away. Declining to guess is the right call; answering with nothing is not.

Added

  • An unknown verdict now hands back the last lines of the log that carried output, under a heading that says exactly what they are: raw log, not a diagnosis. The heuristic is positional, never lexical — it asks where the log stopped, a question with the same answer in every ecosystem, including the ones no rule covers. Optional log_tail key in ci-result.v1 (max 5 lines, each capped at 300 characters and passed through redact_ci_log). (#399)
  • The same tail follows a verdict that is only a hint (under LOW_CONFIDENCE_THRESHOLD = 0.35). rails/rails run 29648807728 is the case: three bundle invocations carry ruby_bundle_failure at 0.3, none of them failed, and the step that actually broke — ./bin/rails assets:precompile — sat in the last lines nobody was shown. Same extraction, same cap, same redaction; only the condition widened. (#400)

failure_class, confidence, signals and exit codes are untouched, and a log that classifies produces byte-identical output to 0.7.5.

Full changelog: https://github.com/patchrail/patchrail/blob/main/CHANGELOG.md

v0.7.5

Choose a tag to compare

@PabloCodes7 PabloCodes7 released this 22 Jul 13:04
6f8bc59

A low-confidence verdict now reads like one.

0.7.4 stopped a class carried by nothing but tool invocations from claiming diagnosis-level confidence, but it printed that capped verdict in exactly the same shape as a confident one — Root cause: <class> / Confidence: 0.3 — so a maintainer still read "root cause" and went off to fix a file that was fine. The number dropped; the message didn't.

Fixed

  • Both human-facing reports (--format text and --format markdown) now add an explicit caveat below a configurable threshold (LOW_CONFIDENCE_THRESHOLD = 0.35): that the verdict is a hint rather than a proven cause, that the signals behind it only show a tool ran and not that it failed, and that the next useful step is the raw log — plus the existing CI failure fixture issue template if the real cause turns out to be something PatchRail cannot classify yet. The guard keys on confidence alone, never on a class or ecosystem; unknown keeps its own existing message instead of being told twice. (#397)

Presentation only: the class, the confidence, the exit code and the --format json payload the GitHub Action consumes are unchanged, and a high-confidence report is byte-identical to 0.7.4.

Full changelog: https://github.com/patchrail/patchrail/blob/main/CHANGELOG.md

v0.7.4

Choose a tag to compare

@PabloCodes7 PabloCodes7 released this 21 Jul 23:01
cf5c08b

Ten false positives removed. Each one is a verdict 0.7.3 gets wrong on a real CI log — the classifier named a tool that was fine, or reported diagnosis-level confidence for a guess.

Fixed

  • A verdict carried by invocations alone now caps at 0.3 confidence. An invocation proves a tool ran, never that it failed; rails/rails run 29648807728 dies on a Ruby SyntaxError and three bundle invocations were producing ruby_bundle_failure at 0.89, sending a maintainer to debug a Gemfile that is fine. Legitimate last resorts (a bare pytest) keep both class and confidence. (#395)
  • A yarn install that finishes with warnings is no longer a broken dependency install, and a git checkout that recovers is no longer a checkout failure. (#393)
  • debconf: (TERM is not set, ...) and other unset terminal/locale variables no longer read as a missing repository secret. (#392)
  • A parenthesized Gateway Timeout (504) is network_transient_failure, not unknown. (#391)
  • mypy listed in a pixi/conda dependency manifest is no longer a type-check failure. (#390)
  • dune's make test and a Crystal spec run are no longer read as C/C++ builds. (#384, #385)
  • Cabal's Could not resolve dependencies: is no longer a Maven build failure. (#383)
  • Flutter's cached Gradle Wrapper is no longer a Gradle build failure. (#382)
  • A passing test's title is no longer a secrets failure. (#380)
  • A failing PHPUnit assertion is no longer a Composer failure. (#378)
  • An installed pytest is no longer reported as a failing test run. (#376)
  • --out writes only to the file, not also to stdout; redact and pilot-pack error on an empty log. (#370, #371)

Changed

  • patchrail --help now ends with a link to the project home.
  • The CI triage job summary shows the local reproduction command, and no longer carries internal adoption telemetry. (#372, #373)
  • PyPI keywords and trove classifiers broadened to match how a CI-triage tool is actually searched for.

Two real mainstream CI logs (spring-boot, kafka) are now pinned as regression tests under tests/data/realworld/, and docs/real-world-benchmark.md covers twenty-one committed public logs.

Full changelog: https://github.com/patchrail/patchrail/blob/main/CHANGELOG.md

v0.7.3

Choose a tag to compare

@PabloCodes7 PabloCodes7 released this 16 Jul 02:34
6210109

Four new-user UX fixes merged since 0.7.2, now on PyPI.

Fixed

  • Piping a non-UTF-8 CI log no longer crashes. The headline one-liner gh run view <run-id> --log-failed | patchrail ci explain feeds raw log bytes to stdin, and real CI logs often carry bytes that aren't clean UTF-8 (a latin-1 accent in a test name, stray ANSI/control bytes, a truncated multibyte sequence). explain, classify, pilot-pack and redact now read the stdin pipe as bytes and decode with errors="replace" — the same forgiveness the --log <file> path already had — instead of raising an uncaught UnicodeDecodeError on a first-timer (#368).
  • ci explain with no --log in a terminal no longer hangs. Running patchrail ci explain without piping anything in used to block forever on stdin.read(), looking frozen. It now detects an interactive terminal and fails fast with a hint to pass --log or pipe a log in, exit 2. A real pipe or file redirect is untouched (#366).
  • --log at a directory or an unreadable file gives a clear message, not a traceback. patchrail ci explain --log logs/ used to leak a raw IsADirectoryError/PermissionError. It now reports the problem on stderr and exits 2, the same clean contract as the missing-file case (#365).
  • A passing CI log is no longer reported as an unrecognized failure to file a fixture for. A green log that matched no failure rule landed on unknown (0.15) with "Open a CI failure fixture issue" — inviting non-failures into the tracker. PatchRail now recognizes a log that plainly announces success, flags likely_successful_run, and replies "No failure detected — point me at the failed run." Conservative: any failure tell vetoes it, so a genuinely unrecognized failure keeps its unknown verdict (#364).

Also includes a GitLab CI clang++ undefined-reference link fixture from @sahilmathur254 (#363).

v0.7.2

Choose a tag to compare

@PabloCodes7 PabloCodes7 released this 15 Jul 08:03
d77de3a

Three fixes merged since 0.7.1, now on PyPI.

Fixed

  • Flyway runtime migration failures are recognised. A flyway migrate step that fails against the database — ERROR: Migration V2__add_users.sql failed, SQL State : 42S01 — used to fall through to unknown (0.15). It now classifies as database_migration_failure at 0.71 and points you at the SQL error to fix; the SQL State : 00000 success code is explicitly excluded so green runs are untouched. Thanks to @hkJerryLeung for the fix (#361, closing #265) — PatchRail's first merged external contribution.
  • A failed yarn run script is no longer reported as a dependency install. yarn classic's docs footer was matched on the bare host, so a yarn run prettier that failed on a diff got sent to reconcile a lockfile. The footer is now pinned to install/add; a real yarn install failure is unaffected (#360).
  • A container OOM-killed under test is no longer reported as the runner running out of memory. OOMKilled / exit 137 / Out of memory from a container runtime's own tests now defer to the concrete cause the log recorded (here go_test_failure). A real host exhaustion still trips a terminal signal and keeps its verdict (#358).

pip install --upgrade patchrail · patchrail --version → patchrail 0.7.2

v0.7.1

Choose a tag to compare

@PabloCodes7 PabloCodes7 released this 14 Jul 22:05
678b10b

Two misreadings, both the same shape: a line that merely names a tool, read as proof that tool had failed.

Fixed

  • A job that declares timeout-minutes is no longer reported as a job that timed out. Actions echoes your step config into the log — timeout-minutes: 180 — on green runs and red ones alike, and PatchRail was reading that declaration as evidence the limit had been hit. On envoyproxy/envoy's coverage run 29363920524 the job never came near its 180-minute ceiling: one directory out of 430 had slipped three tenths of a point under its coverage threshold. That log now answers code_coverage_threshold. A job that really did run long is unaffected — the runner says so in words (has exceeded the maximum execution time of…, The operation was canceled), and those still land as ci_job_timeout.

  • A Go job that fails to compile is no longer reported as a lint failure. golangci-lint-action names its tool five times just to ship it — version lookup, cache hit, install, echoed command line, timing — and PatchRail was reading those install lines as proof the linter had failed. Any Go job that merely declared the action got a confident go_lint verdict, whatever had actually broken. On grafana/grafana's lint-go run 27635190952 there was no lint finding at all: a function had grown a parameter and eight call sites in one test file had not. That log now answers go_test_failure and reproduces with go test ./.... A linter that really did report something ((gofmt), (gci), (revive)) is still go_lint.

  • Go call-site mismatches are recognised. not enough arguments in call to and too many arguments in call to — the compiler's own words — no longer pass unread, alongside undefined:.

Both fixes were measured against the real failing logs, which ship in the repo under examples/real-world/ so you can reproduce the before and after yourself.

pip install --upgrade patchrail
patchrail ci explain --log examples/real-world/envoy-29363920524-excerpt.log

Full changelog: https://github.com/patchrail/patchrail/blob/main/CHANGELOG.md

v0.7.0 — nine fixes for one mistake

Choose a tag to compare

@PabloCodes7 PabloCodes7 released this 14 Jul 20:33
29db167

Nine misreadings, all of one kind.

Every fix in this release started as a real failing run at a real project — pandas, deno, svelte, istio, astro, ruff, prometheus, pytorch — where PatchRail answered confidently and answered wrong. And every one of them was wrong in the same way: it mistook a line that merely named a tool for a line proving that tool had failed.

An install listing. A script echoed before it runs. A suppressible CMake warning. A cache cleaning up after a job that was already dead. A test quoting a compiler back at itself.

A verdict that no signal ever witnessed failing now yields to what the runner actually flagged — and when nothing is left, the honest answer is unknown, with the failing line handed back to you.

What stops being wrong

  • A CMake policy that "is not set" is no longer read as a repository secret that is not set (CMP0148 is shaped exactly like an environment variable). pytorch's lint job died in jq; we sent maintainers to audit their secrets.
  • A dependency Dependabot merely listed in its job definition is no longer mistaken for a linter that ran and failed. svelte's maintainers were told pnpm lint broke. No linter ran.
  • Output a test quotes back at you is no longer read as the job's own diagnostic. deno's spec suite typechecks deliberately-broken programs on purpose; we told the maintainers of a TypeScript runtime that their TypeScript was broken, at 0.95 confidence.
  • The source of a step, echoed before it runs, is no longer read as the step's output. ruff's Rust panic came out python_test_failure, on an error branch for a download that had succeeded.
  • A cache that could not save, and a tool the job merely installed, no longer decide the verdict. pandas' doc build sent a maintainer to debug a cache that was working perfectly.
  • A failed npm audit is reported as a failed scan, not a broken install — and npm's post-install audit tally is not a failed scan at all.
  • A tie between two failure classes goes to the rule that watched something fail. prometheus — a Go repo, whose Go tests failed — came out javascript_lint.
  • A proxy logging its own client disconnects is not a network outage.

New: the classifier, measured against real logs

docs/real-world-benchmark.md runs PatchRail against seven real CI logs and shows what it said before, what it says now, and the command that reproduces each one. The logs are committed verbatim under examples/real-world/, so you can check the numbers instead of trusting them.

The misses are in the table, not in a footnote. pandas still lands on python_test_failure at 0.53 rather than naming the crash. grafana calls a compile error go_lint. Four of the seven verdicts are identical before and after — which is the point: these fixes are narrow, and a benchmark that only shows wins is not a benchmark.

pipx install patchrail
patchrail ci explain --log failed-ci.log

Runs locally. No API keys, no log upload.

v0.6.1

Choose a tag to compare

@PabloCodes7 PabloCodes7 released this 14 Jul 13:24
e3ff1bf

Patch release: a success announced through the error channel is no longer reported as somewhere to "start".

oven-sh/bun's failing run carries exactly one runner annotation in 4,709 lines, and the workflow emits a success through it — ##[error]✅ Autofix task started. The unknown verdict handed that line back under "the CI runner did annotate these lines as errors — start there", pointing the maintainer at a line saying everything went fine.

The verdict itself was, and stays, correct: unknown is honest for that log. Only the evidence changes. An annotation is now dropped when it opens with a success mark and names no failure anywhere in the line — a guard deliberately lopsided towards keeping, so ✅ 2 passed, ❌ 1 failed and ✔ image built, but the upload failed both survive.

Users of the CI Triage Action pick this up on their next run with no change on their side.

Closes #329. Full notes in CHANGELOG.md.

v0.6.0 — an unknown log that explains itself, and no more invented diagnoses

Choose a tag to compare

@PabloCodes7 PabloCodes7 released this 14 Jul 12:34
9aadef6

Three changes that landed after 0.5.0, all found the same way: piping real failing runs from public repositories through the one-liner in the README, the way a first-time user meets this tool. Eight new ecosystems (deno, bun, istio, astro, airflow, envoy, ruff, pydantic) — three came back confidently wrong.

Fixed

  • A tool named inside a filename is no longer a diagnosis. A CI log in a monorepo is mostly filenames, and a rule could be carried start to finish by a word that only ever appeared inside one. oven-sh/bun — a Zig/JS runtime — was told its pnpm lockfile was out of date, on the evidence of prettier listing files it left (unchanged), two of which sat under a lockfile/ directory; in the same log, a regression test named after an advisory (GHSA-pfwx-36v6-832x.test.ts) was read as a failed security scan that never ran. istio/istio, a Go repository, was diagnosed with both a Java build failure and a Node install failure, read out of a 2929-character single-line Dependabot JSON blob. withastro/astro matched on --no-frozen-lockfile — the flag that permits the change it was complaining about. A signal must now match outside a path token to carry a verdict, and a bare tool name corroborates a failure without ever constituting one.
  • The runner's content-free annotations no longer dress an empty answer up as a finding. Process completed with exit code 1. appears in every failing run, whatever the cause; reported back under a heading it read as evidence, and — each exit code being a distinct string — a matrix build's worth of them would crowd out the annotation that actually names the failure.

Added

  • An unknown verdict now hands back the line the runner itself flagged. A log no rule matches usually still names its own failure: psf/requests failed on a single, self-explanatory line — ##[error]"github-token" length must be less than or equal to 100 characters long — and PatchRail answered unknown, no signals, nothing else. Those annotations are now reported in runner_errors (JSON), under "Errors the runner reported" (Markdown), and as Runner reported: (text), redacted and de-duplicated. They stay evidence and nothing more: an annotation says where the job died, not why, so the class stays unknown and the confidence stays at 0.15. PatchRail still does not pretend to recognize a log it does not recognize.

Genuine failures are unaffected: real lockfile, security-scan and build failures score exactly as before, and no fixture changes classification or confidence (221 fixtures, top-1 1.0).

pip install --upgrade patchrail

v0.5.0 — real logs, real causes

Choose a tag to compare

@PabloCodes7 PabloCodes7 released this 14 Jul 10:32
4e063b2

Five accuracy fixes that landed after 0.4.0, all found by running ci explain over real failing runs from public repositories rather than over the fixture zoo. The zoo's logs are clean; a real one is not, and every one of these bugs hid in the difference.

Fixed

  • A tool the job merely named no longer outranks the one that failed. set -x echoes every command into the log whether it passes or not, and pip announces every dependency it resolves — so Collecting mypy==1.17.1 put mypy in the log of a job that never type-checked anything. Signals are now judged by the line they land on. Caught 3 of 9 real Python runs (httpx, prefect, home-assistant) and 3 of 17 across Go/Rust/JVM/.NET/C++ (kafka, moby, opencv), where a GRADLE_HOME= entry in a Windows environment table was enough to diagnose a Java build failure in a job that compiled nothing.
  • ANSI colour codes no longer hide the failure. CI keeps the colour on and the GitHub log API serves the escapes back, so the reset lands inside the failure line (^[[31mFAILED^[[0m tests/…) and FAILED .*:: silently stopped matching. A real airflow test failure was reported as an artifact error.
  • python_test_failure now recognises pytest's own verdict — 1 failed, 1416 passed in 18.37s, the short test summary, and collection errors, not just named tests.
  • Post-failure cleanup noise no longer outranks the real cause. The if: failure() step that uploads logs and finds nothing is the commonest shape in CI; it used to tie the real cause and win on declaration order.
  • The classifier no longer reads a command that merely ran as the failure — \btsc\b matched the x86 time stamp counter in /proc/cpuinfo, diagnosing a rust-lang/rust build failure as a TypeScript typecheck.

Genuine failures are unaffected: a real gradle, clippy, docker or dotnet failure scores exactly as before, and no fixture changes classification or confidence (221 fixtures, top-1 1.0).

Added

  • patchrail schema ci-classes publishes the schema for ci classes --format json, with a conformance test that pins the emitted schema_version to the schema you can fetch, so the contract cannot drift silently again.
  • docs/json-cookbook.md documents the ci classes payload with a coverage recipe.
pip install --upgrade patchrail