Skip to content

fix: declare the supertool dependency as supertool-cli (#1806) #2071

fix: declare the supertool dependency as supertool-cli (#1806)

fix: declare the supertool dependency as supertool-cli (#1806) #2071

Workflow file for this run

name: tests
on:
push:
branches: [main]
pull_request:
# A dropped push run (#679) has no remedy short of an empty commit to main without
# this: `gh workflow run tests.yml` is refused for a workflow with no
# workflow_dispatch, and `gh run rerun` needs a run id that does not exist when
# nothing was ever created. Deliberately not on changelog.yml -- see
# tests/test_workflow_dispatch_679.py for why that would be actively broken rather
# than merely unused. `pytest` now reads `github.event_name` and
# `github.event.inputs.full_matrix` (#1246, below) to size its matrix, so a bare
# dispatch (input left at its `false` default) still behaves like a push run --
# only `full_matrix: true` diverges from that.
# #1246: `full_matrix` defaults false, so an ordinary dispatch (the #679
# dropped-push-run remedy) still mirrors the reduced push matrix below rather than
# silently ballooning to every OS x Python combination. The release process sets it
# true when it dispatches this workflow against the commit about to be tagged, which
# is the only route to the full matrix -- see the `exclude:` expression on the
# `pytest` job's matrix.
workflow_dispatch:
inputs:
full_matrix:
description: "Run every OS x Python leg instead of the reduced push/PR set (release gate use)"
required: false
default: false
type: boolean
# One run per ref at a time, and the superseded one is cancelled -- but only for a pull
# request (#962). Three pushes to one branch in fifteen minutes used to start 42 legs and
# leave 28 of them running against commits nobody would merge; worse than the minutes, a
# pull request's summed check tally mixes legs from a superseded commit with legs from the
# head commit, so "not all green" can mean an old run that will never finish.
#
# `cancel-in-progress` is an expression rather than a bare `true` because a push run on the
# default branch must never be cancelled: it is what `scripts/statusline.py` and the release
# gates read as that commit's verdict, and cancelling one leaves the commit with a run that
# reports neither pass nor fail -- an absence produced by our own configuration, read later
# as an absence of findings. `github.ref` already differs between the two event types
# (`refs/pull/N/merge` vs `refs/heads/main`), so they could not share a group in any case;
# the condition states the intent where a reader of this file will look for it, rather than
# leaving it as a property of GitHub's ref naming that a future group change could lose.
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
# Workflow level: every job here runs tests or lints and reports, none writes to the
# forge, and a job added later inherits read-only instead of the repository default of
# read/write (#32). Three jobs as of #635 (pytest, shell, lint) -- this line names none
# of them by count on purpose, so a fourth job does not make it stale again.
permissions:
contents: read
jobs:
pytest:
runs-on: ${{ matrix.os }}
# #1658: 15 left no headroom on windows-latest, where the suite itself finishes
# (green) in ~862-873s (0:14:22-0:14:33) and checkout/setup/dependency-install
# overhead pushes the job's total wall time past the ceiling -- three
# back-to-back cancellations on the same green commit (jobs #105401634166,
# #105406249723, #105410692355). 20 restores several minutes of real margin over
# the slowest observed leg without removing the cap on a genuine hang.
#
# #1660: 20 was itself cancelled again, twice, on PR #1654's own commit
# `fad72a25` (jobs #105421966394, #105427076121) -- the suite's own SELF-
# REPORTED runtime (its own "N passed... in Xs" line) had grown from ~873s to
# ~1176.55s (0:19:36), leaving only ~24s of job-wall-clock margin against the
# 20-minute cap. Two things this file does NOT claim, having been caught
# overclaiming both in review: whether that ~300s of growth is itself test
# count/content growth, runner variance, or a genuine slowdown is unresolved
# (open question, not "not more tests running"); and how long the process was
# actually still alive after its own summary line printed is UNKNOWN, not
# "several minutes" -- both runs were cut off by the cap itself, so the true
# size of any post-summary gap was never observed, only that it is non-zero
# (in both runs a `KeyboardInterrupt` lands in the main thread, blocked in a
# `threading` wait/join, after the summary already reported every test
# passing). Root cause not identified; #1660's own issue text asks for exactly
# the diagnostic `tests/posthang_diagnostics_1660.py` now adds: periodic
# thread-stack dumps once a session has finished but the process has not
# exited. 30 is deliberately generous rather than a repeat of the same narrow
# margin #1658 used, since the runtime has already grown once, unexplained,
# and may again before the diagnostic pins down why.
#
# It did grow again, and the diagnostic did not help: a THIRD occurrence
# (job #105442284149, commit `c201a8a0`, carrying #1660's own faulthandler
# diagnostic) still ended `cancelled`, `1773.10s (0:29:33)` against this
# 30-minute cap -- only ~26.9s of margin, and no `faulthandler` output
# anywhere in that job's log. That is not the diagnostic failing to catch a
# hang; it never got the chance to -- `POST_SESSION_DUMP_AFTER_SECONDS` was
# 60 at the time, so the first dump was scheduled for 33s AFTER this cap
# would already have killed the job. Fixed in
# `tests/posthang_diagnostics_1660.py` (lowered to 15, with a coupling test
# in `tests/test_posthang_diagnostics_1660.py` that catches THIS delay
# going stale against `tests/test_pytest_leg_timeout_1658.py`'s own
# hand-maintained runtime constant -- not a live measurement, so a FOURTH
# occurrence still needs that constant updated by hand before the coupling
# test can say anything about it). This cap is left at 30 rather than
# raised a third time: the growth (873s -> 1176.55s -> 1773.10s) has tracked
# each cap raise closely enough that raising it again is not expected to
# hold either, and #1660's own issue text argues against repeating that fix
# for however long the trend continues (this was true when this cap was
# last left at 30; #1784 below raises it a third time anyway, but for a
# different, now-measured reason -- see that comment). The next
# recurrence, if the diagnostic can now actually fire, is what settles
# whether this is a real hang or something else growing the suite's own
# reported runtime.
#
# #1784: this 30-minute figure was clearing OBSERVED_WORST_CASE_SUITE_MINUTES
# (`tests/test_pytest_leg_timeout_1658.py`) by only ~26.9s, and that whole
# margin was computed against pytest's own self-reported runtime alone --
# `timeout-minutes` counts from JOB start, so checkout/setup-python/the
# Windows Defender exclusion step/`pip install`, none of which pytest's own
# clock ever sees, were silently eating into it. Measured directly from the
# GitHub Actions API for the six jobs already cited above and in
# `tests/posthang_diagnostics_1660.py`, that pre-pytest overhead ran 26-43s
# -- already comparable to, and at its worst exceeding, the entire 26.9s
# margin the arithmetic assumed was free. Raised by 2 minutes specifically
# to cover that now-measured, largely fixed overhead (not to chase the
# suite's own separate, unexplained runtime growth the paragraph above
# already argues against re-raising for) -- see
# `tests/test_pytest_leg_timeout_1658.py`'s own `PRE_PYTEST_OVERHEAD_SECONDS`
# for the six measured values this raise is sized against.
timeout-minutes: 32
# #1246: the `os` x `python-version` product below is the full 3x4 release matrix,
# unconditionally -- so `full_matrix: true` restores full coverage by excluding
# nothing, rather than needing a second, separately-maintained matrix to drift out
# of sync with this one. An ordinary push/pull_request, or a workflow_dispatch left
# at its `false` default, drops to the five legs #1246 measured and kept: ubuntu
# 3.12 (cheap default), ubuntu 3.9 (the declared floor), windows 3.12 (the observed
# failure platform), macos 3.12 (the third OS), plus the unconditional `shell` job
# below. `exclude:` lists the eight combinations that reduction drops; each entry
# must name both `os` and `python-version` or GitHub Actions treats it as a
# wildcard on the field left out, excluding every python-version for that os (or
# vice versa) instead of the one combination intended.
strategy:
fail-fast: false
matrix:
os: [ubuntu-latest, macos-latest, windows-latest]
python-version: ["3.9", "3.10", "3.11", "3.12"]
exclude: ${{ (github.event_name == 'workflow_dispatch' && github.event.inputs.full_matrix == 'true') && fromJson('[]') || fromJson('[{"os":"ubuntu-latest","python-version":"3.10"},{"os":"ubuntu-latest","python-version":"3.11"},{"os":"macos-latest","python-version":"3.9"},{"os":"macos-latest","python-version":"3.10"},{"os":"macos-latest","python-version":"3.11"},{"os":"windows-latest","python-version":"3.9"},{"os":"windows-latest","python-version":"3.10"},{"os":"windows-latest","python-version":"3.11"}]') }}
env:
# The exact CPython patch `setup-python` resolves "3.9" to on the hosted
# windows-latest runner today. Both the cache path and the cache key below are
# derived from it, and the disclosure step checks the interpreter that actually
# ran against it -- two literals that must agree is a bug waiting for the day
# one of them is edited without the other.
WIN_PY39_PATCH: "3.9.13"
steps:
- uses: actions/checkout@v7
# CPython 3.9 has left the windows-latest image, so `setup-python` reports
# `Version 3.9 was not found in the local cache` and downloads and installs it --
# measured at 44s of a 229s critical-path leg under xdist (#1196; #1177 landed
# first and made this leg the critical path in the first place). Restoring the
# hosted tool cache here is the whole fix: `setup-python` then finds the tree
# and skips the download.
#
# `Digital-Process-Tools/claude-supertool` already ships this pattern for its
# own CI (#1127 there); this step and its trap comment are ported from that
# workflow (confirmed via `gh api
# repos/Digital-Process-Tools/claude-supertool/contents/.github/workflows/tests.yml`
# rather than retyped from a description).
#
# The cached path is the VERSION directory, not `.../x64` beneath it.
# `@actions/tool-cache` records an installed tool by writing a sibling marker
# `x64.complete` next to `x64/`, and its `find()` returns a miss when that
# marker is absent -- so caching only `x64/` restores a tree `setup-python`
# then ignores, and every run is a cache HIT that saves nothing.
#
# Windows-only and 3.9-only on purpose: the other three Windows legs already
# have the interpreter baked into the image, and macOS pays a similar ~22s for
# 3.9 but is never the critical path (107s against Windows' 266-305s), so
# caching there would buy billed minutes and no wall clock.
#
# A cache written on a pull request branch is visible only to that branch: the
# `push: [main]` run is what populates the key every other branch restores
# from, so the first PR after this change lands may still miss -- expected, not
# a failure. The "Tool cache state" step below discloses which happened rather
# than leaving it to be inferred from `Set up Python`'s own duration.
- name: Restore the CPython ${{ env.WIN_PY39_PATCH }} tool cache (#1196)
id: toolcache-39
if: runner.os == 'Windows' && matrix.python-version == '3.9'
uses: actions/cache@v6
with:
path: ${{ runner.tool_cache }}/Python/${{ env.WIN_PY39_PATCH }}
key: ${{ runner.os }}-${{ runner.arch }}-toolcache-python-${{ env.WIN_PY39_PATCH }}
- name: Set up Python ${{ matrix.python-version }}
uses: actions/setup-python@v7
with:
python-version: ${{ matrix.python-version }}
# Windows Defender's real-time scanner walks every file this job touches on the
# runner's own disk, and the suite creates many small temp files per assertion
# (each `mktemp` fixture, each rewritten .claude/jit-context/ under test). This is
# a documented GitHub Actions cost specific to windows-latest: macOS and Linux
# runners carry no equivalent always-on scanner. Excluding the checkout and the
# runner temp directory removes that tax without touching what any test asserts --
# it changes nothing about correctness, only how fast the disk answers. Ported
# from Digital-Process-Tools/claude-jit-context#310 (open, not yet merged there,
# at the time this landed -- the shape below is that PR's actual diff, not a
# retyped guess). The measured saving on this repo's own suite is pending a live
# CI round; see #938.
#
# Each Add-MpPreference call is wrapped in try/catch rather than given a bare
# continue-on-error (#1081): the cmdlet has been observed to fail on a hosted
# runner with 0x800106ba (the Defender service itself unavailable), a cause
# unrelated to the change under test, and an unguarded step failure there fails
# the whole leg and renders identically to a real Windows test failure in
# gh-pr:N:status -- distinguishable only by opening the job log.
# continue-on-error alone would stop that but stays silent about which
# happened; try/catch keeps the step from failing the job AND lets one place
# disclose "applied" vs "failed" per path, reusing the hit/miss/skip
# disclosure shape #1196 already put in this workflow (see "Tool cache state"
# below) rather than inventing a second one.
- name: Exclude the checkout and temp dir from Windows Defender scanning
if: runner.os == 'Windows'
shell: pwsh
run: |
foreach ($path in @("${{ github.workspace }}", "$env:RUNNER_TEMP")) {
try {
Add-MpPreference -ExclusionPath $path -ErrorAction Stop
Write-Host "defender-exclusion: ok ($path excluded)"
} catch {
Write-Host "::warning::defender-exclusion: failed ($path not excluded -- $($_.Exception.Message))"
}
}
# Three states, never a silent one. `cache-hit: true` on its own is not
# evidence the cache did anything: `WIN_PY39_PATCH` is a literal, and if
# `setup-python` ever resolves "3.9" to a different patch, the step above would
# restore -- and then re-save -- a directory nothing reads. A permanent HIT
# sitting next to a permanent 44s download is this repo's own defect class
# exactly (an absence produced by the tool, read as an absence in the world),
# so the check is against the interpreter that actually ran, not against the
# cache action's own opinion of itself.
#
# `::error::` annotates the run summary without failing the step: a
# runner-image change is not the fault of whoever pushes next.
#
# NOT `always()`: GitHub runs `bash` with `set -eo pipefail`, so if the step
# above failed there is no guarantee `python` behaves the same way, and a step
# that cannot run is not itself a finding -- it is reported as `skipped`.
- name: Tool cache state (#1196)
if: runner.os == 'Windows'
shell: bash
env:
PY_VERSION: ${{ matrix.python-version }}
CACHE_HIT: ${{ steps.toolcache-39.outputs.cache-hit }}
run: |
if [ "$PY_VERSION" != "3.9" ]; then
echo "toolcache : skipped (CPython $PY_VERSION is baked into the windows-latest image; only 3.9 is missing from it)"
exit 0
fi
resolved=$(python -c "import sys; print('.'.join(map(str, sys.version_info[:3])))" || true)
if [ -z "$resolved" ]; then
echo "toolcache : skipped (could not read the interpreter version -- this says nothing about the cache either way)"
exit 0
fi
if [ "$resolved" != "$WIN_PY39_PATCH" ]; then
echo "::error::toolcache stale pin - this workflow caches CPython $WIN_PY39_PATCH but setup-python resolved $resolved, so the restore can never be read. Set WIN_PY39_PATCH to $resolved in .github/workflows/tests.yml."
echo "toolcache : finding (stale pin: cached $WIN_PY39_PATCH, running $resolved -- the restore is dead weight)"
exit 0
fi
if [ "$CACHE_HIT" = "true" ]; then
echo "toolcache : ok (hit -- CPython $resolved restored, no download)"
else
echo "toolcache : miss (CPython $resolved downloaded; the tool cache is saved at job end, for the next run -- expected on the first run after this change on a given branch, since a cache written on a PR branch is visible only to that branch)"
fi
# markdown-it-py is not a test tool, it is the assembler's one dependency, and the
# suite over scripts/assemble_changelog.py cannot reach a single assertion without
# it -- the script refuses to run rather than fall back to text scanning. Absent
# here, that whole suite errored on every leg, so the heading refusal, the raw-HTML
# refusal and the fence state machine had never been exercised on a runner at all;
# they passed locally because the package happened to be installed there. The
# changelog workflow learned this the same way and the lesson was applied to that
# one file only. tests/test_workflow_dependencies.py now holds both against
# scaffold.ASSEMBLER_DEPENDENCIES, which is where the package name is declared.
# pyyaml is here for tests/test_shell_leg_budget_303.py, which asserts against the
# parsed `shell` job rather than against substrings of this file -- a regex over a
# workflow is the shape that keeps passing while the workflow is broken. It is the
# only test that needs it; the rest of the suite reads workflows as text on
# purpose. That file fails rather than skips when `CI` is set and pyyaml is
# missing, so dropping this package cannot make it go quiet.
#
# ruff is here (#635) so tests/test_ruff_ratchet_635.py actually runs instead of
# skipping: that file's own `pytestmark` declines every test when `ruff` is not on
# PATH, and this is the only job that runs `pytest` at all -- the `lint` job below
# never does. Absent here, the ratchet script's own logic (baseline comparison,
# exit-code selection, the third `could-not-run` state) would ship a whole new file
# of tests that pass by skipping on all 12 legs of this matrix, every time.
- name: Install test deps
run: |
python -m pip install --upgrade pip
pip install pytest pytest-cov pytest-xdist markdown-it-py pyyaml ruff==0.16.3
# #1177: what `pytest -n auto` would request on THIS leg, from doctor.py's own
# transcription of xdist's worker-count logic (#367), before any decision is taken
# about parallelising this matrix. The runners' core counts have never been read
# here -- the 2-3x figure in #1177 is reasoned from a documentation page, and
# GitHub has changed those numbers before (public repositories went from 2 to 4
# vCPU in 2024). Twelve legs printing their own answer is the reading that
# argument needs, and it is worth keeping afterwards: the number a later
# regression would move is then in every run's log rather than in one issue.
#
# doctor.py exits 0 in every mode, so this can never fail a leg, and it reports
# `worker sizing unknown` rather than a number when nothing on the runner
# answered. Unconditional on purpose: a leg that skips it reports nothing, and
# the platforms this matrix exists for are exactly the ones whose sizing is not
# known here. psutil is deliberately not installed for it -- xdist would not find
# one either, so the source this prints is the source `-n auto` would really use.
- name: Report what `pytest -n auto` would request here
run: python scripts/doctor.py --worker-sizing
# #1177 candidate 1: `-n auto --dist loadfile`. The probe added by #1189 has now
# reported on all three platforms of this matrix -- ubuntu 4 workers, windows 4,
# macOS 3, ubuntu's read from three legs of run 34056088247 -- so the parallelism
# available here is measured rather than taken from a documentation page.
#
# `--dist loadfile` and not the default `--dist load`: this suite's tests build git
# repositories, change working directory and manage worktrees, and keeping one
# file's tests on one worker is what lets those fixtures survive being run beside
# each other. It is the weaker parallelism and it is the one worth trying first.
#
# What this does NOT buy, so nobody reads the worker count as a speedup factor: a
# whole file lands on one worker, so the longest single FILE sets the floor.
# `tests/test_no_test_pins_the_current_version_350.py` alone was 25.33s of ubuntu
# 3.12's 234.23s of measured test time, and no number of workers divides that.
#
# The risk here is not wall clock, it is a suite that passes in a different order.
# A green first run is the outcome to distrust rather than the one to ship: a race
# that fires one run in ten arrives later as a flake nobody attaches to this change.
#
# What xdist did NOT cause, recorded because this change was blamed for it for two
# rounds: the 123 launcher-suite failures on the Windows legs came from the
# `shell: bash` the coverage step briefly carried, and reproduced on a serial run
# with no workers in the log at all. See that step's own comment.
#
# #1176: 126 of the 428 tracked test files counted at filing time are content
# guards over this repo's own markdown/config -- their answer does not vary by OS
# or interpreter, and running
# them on all 12 legs buys nothing a reader ever sees. Marking every one of them
# was weighed and declined: the naive "no tmp_path/subprocess/monkeypatch" heuristic
# selects the CHEAP tests (9.67s total, ran locally), not the expensive ones -- the
# nine tests actually worth deselecting spawn a subprocess or read this repo's whole
# tree, and #1176 marks exactly those nine `invariant` (tests/test_invariant_
# marker_1176.py holds the registry). `-m "not invariant"` on every leg but the full
# one deselects them; the full leg is the SAME leg #1177 already designated to keep
# coverage, `(ubuntu-latest, 3.12)`, so this reuses one special-leg concept instead
# of inventing a second, and it keeps the deselect off the 51.6%-of-wall-clock
# Windows legs.
#
# Both flags live on the same unconditional step, each behind its own `${{ }}`
# ternary with the identical condition -- never two `if:`-guarded steps, because
# two conditions that are both false is a leg that runs no tests and reports
# success anyway (this repository's own defect class, installed in the merge gate
# on purpose). Each ternary is written `X && 'flag' || ''` rather than
# `X && '' || 'flag'`: the empty string is falsy in a GitHub expression, so the
# non-empty arm has to be the TRUE arm or the flag lands on every leg instead of
# eleven of them -- #1177's own fix for the identical mistake on the coverage flag.
- name: Run tests
run: |
# #1673: any single test past 180s dumps its own stack and name to
# stderr (pytest's built-in faulthandler plugin), so a hung test is
# named in the log rather than showing as a bare KeyboardInterrupt
# at the job's own cap -- which four issues read as a post-session
# hang before this named the real one on its first run.
pytest -n auto --dist loadfile -o faulthandler_timeout=180 ${{ (matrix.os != 'ubuntu-latest' || matrix.python-version != '3.12') && '--no-cov' || '' }} ${{ (matrix.os != 'ubuntu-latest' || matrix.python-version != '3.12') && '-m "not invariant"' || '' }}
shell:
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@v7
# Both guards below scoped themselves with `git ls-files '*.sh'`, which returns
# exactly one path here -- scripts/doctor.sh. bin/oss-workspace is tracked, is
# POSIX sh and has no extension, so neither `bash -n` nor `shellcheck` had ever
# read the plugin's own entry point, on any leg, on any platform, in any release,
# and this leg was green throughout: a lint that ran and found nothing and a lint
# that never received the file both exit 0 (#193).
#
# The list is derived by extension OR shebang, in scripts/shell_sources.py, and
# the reasoning for deriving it rather than naming the file here is in that
# script's docstring. Two properties this job depends on:
#
# * it exits 2 when it matches nothing, so an empty selection fails this leg
# instead of running `shellcheck` with no arguments and exiting 0 -- the half
# of the repair that would otherwise reintroduce the bug;
# * it exits 3 when a tracked file could not be read, because `could not look`
# is not `looked and found no shebang`.
#
# Redirection and not a pipe: the default shell here is `bash -e` without
# `pipefail`, so `... | tee list` would report tee's exit status and swallow both
# refusals.
- name: Enumerate shell sources
run: |
python3 scripts/shell_sources.py --root . > shell-sources.txt
echo "linting:"; cat shell-sources.txt
# `while ... done < file` and not `... | while`: a pipeline puts the loop in a
# subshell, and a `while` returns the status of the loop's completion rather than
# of its body, so a failing guard would be swallowed. The status is therefore
# carried out by hand -- `|| exit 1` below, and a flag in the shellcheck step so
# every file is reported rather than only the ones before the first failure.
#
# `xargs -d` would be shorter and is a GNU extension; this form is the one that
# can be run and watched fail on a maintainer's own machine, which is how the
# `| while` trap above was caught rather than reasoned about.
- name: Syntax-check every shell source
run: |
while IFS= read -r f; do bash -n "$f" || exit 1; done < shell-sources.txt
# No package fetch here, and that is the whole of #303 rather than a tidy-up.
#
# This step used to open with `sudo apt-get update -qq && sudo apt-get install -y
# shellcheck`. Measured on job 96152482222 (run 32276977038, main at 7ea64c9) the
# job ran 10.32 minutes against `timeout-minutes: 10` and was killed. GitHub
# renders a timeout kill as `cancelled`, not `failure`, so main and every open
# pull request read as "0 failed, 1 cancelled -- not green" with nothing broken,
# six times in one day.
#
# Two readings decided this rather than a bigger cap:
#
# * the killed step logged nothing at all between its own `##[endgroup]` and
# `##[error]The operation was canceled.` 10m12s later. `apt-get install` is
# not quiet, so it was never reached: the entire window was `apt-get update`;
# * every successful run logged `shellcheck is already the newest version
# (0.9.0-1)`. The install has never installed anything -- shellcheck ships in
# the ubuntu-latest image -- while `apt-get update` cost 68.4s of a 70s step
# on 2026-08-19 and 5.2s on 2026-08-16. The linting itself is under a second.
#
# So the fetch is removed rather than pinned. A pinned release tarball or a cached
# apt layer would each be a network round trip bought for a binary already on the
# runner, and pinning one means carrying a version and a checksum for it.
#
# `timeout-minutes` is deliberately left at 10 and is now a bound on a hang rather
# than a budget this job routinely spends: the measured job is a few seconds. It
# is not raised, because the number it was failing against was almost entirely
# `apt-get`; it is not lowered, because nothing measured argues for a particular
# smaller number and a tight cap reintroduces reds caused by nothing in a diff.
#
# The version now floats with the runner image. That is the cost, and it is made
# visible rather than assumed: the version is printed, and the binary being absent
# is its own exit code. Before this, an absent shellcheck exited 127 per file into
# the same `fail` flag a real finding uses, so `could not lint` and `linted and
# found a problem` reached the leg's status identically -- and `A && B` under
# `bash -e` does not exit on A, so a failed install fell through to exactly that.
- name: shellcheck
run: |
if ! command -v shellcheck > /dev/null 2>&1; then
echo "shellcheck is not on PATH." >&2
echo "It ships in the ubuntu-latest image and this job deliberately fetches" >&2
echo "nothing (#303). If the image has stopped carrying it, install it here" >&2
echo "again -- but pin a version and a checksum, and put the fetch in a step" >&2
echo "of its own so a slow mirror is not charged to this lint's budget." >&2
exit 4
fi
shellcheck --version
fail=0
while IFS= read -r f; do shellcheck -S warning "$f" || fail=1; done < shell-sources.txt
exit "$fail"
# #635 -- 250 tracked .py files had no linter anywhere. The write-time half
# (supertool's `ruff` validator, `.supertool.json`, catches a finding the
# moment a file is edited) cannot see a file nobody touches in this run, and
# is not wired into CI at all, so this leg is what covers the whole tree and
# cannot be skipped. Ubuntu-only and its own job (not folded into `pytest`
# above): the ruleset in pyproject.toml's `[tool.ruff.lint]` does not vary
# by OS or interpreter, so running it 12 times would be 11 wasted runs
# buying no new coverage -- the same reasoning `shell` above already uses.
#
# `scripts/ruff_ratchet.py` is a ratchet, not a "must be zero" gate: #635's
# own measurement found 95 pre-existing findings across 51 files outside
# that issue's claimed scope, so this only fails a PR that makes the count
# go UP, never one that leaves it flat. See that script's own docstring.
lint:
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@v7
- name: Set up Python
uses: actions/setup-python@v7
with:
python-version: "3.12"
- name: Install ruff
run: |
python -m pip install --upgrade pip
pip install ruff==0.16.3
- name: ruff (#635 ratchet)
run: python3 scripts/ruff_ratchet.py --root .