Repository navigation
fix: declare the supertool dependency as supertool-cli (#1806) #2071
Workflow file for this run
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| name: tests | |
| on: | |
| push: | |
| branches: [main] | |
| pull_request: | |
| # A dropped push run (#679) has no remedy short of an empty commit to main without | |
| # this: `gh workflow run tests.yml` is refused for a workflow with no | |
| # workflow_dispatch, and `gh run rerun` needs a run id that does not exist when | |
| # nothing was ever created. Deliberately not on changelog.yml -- see | |
| # tests/test_workflow_dispatch_679.py for why that would be actively broken rather | |
| # than merely unused. `pytest` now reads `github.event_name` and | |
| # `github.event.inputs.full_matrix` (#1246, below) to size its matrix, so a bare | |
| # dispatch (input left at its `false` default) still behaves like a push run -- | |
| # only `full_matrix: true` diverges from that. | |
| # #1246: `full_matrix` defaults false, so an ordinary dispatch (the #679 | |
| # dropped-push-run remedy) still mirrors the reduced push matrix below rather than | |
| # silently ballooning to every OS x Python combination. The release process sets it | |
| # true when it dispatches this workflow against the commit about to be tagged, which | |
| # is the only route to the full matrix -- see the `exclude:` expression on the | |
| # `pytest` job's matrix. | |
| workflow_dispatch: | |
| inputs: | |
| full_matrix: | |
| description: "Run every OS x Python leg instead of the reduced push/PR set (release gate use)" | |
| required: false | |
| default: false | |
| type: boolean | |
| # One run per ref at a time, and the superseded one is cancelled -- but only for a pull | |
| # request (#962). Three pushes to one branch in fifteen minutes used to start 42 legs and | |
| # leave 28 of them running against commits nobody would merge; worse than the minutes, a | |
| # pull request's summed check tally mixes legs from a superseded commit with legs from the | |
| # head commit, so "not all green" can mean an old run that will never finish. | |
| # | |
| # `cancel-in-progress` is an expression rather than a bare `true` because a push run on the | |
| # default branch must never be cancelled: it is what `scripts/statusline.py` and the release | |
| # gates read as that commit's verdict, and cancelling one leaves the commit with a run that | |
| # reports neither pass nor fail -- an absence produced by our own configuration, read later | |
| # as an absence of findings. `github.ref` already differs between the two event types | |
| # (`refs/pull/N/merge` vs `refs/heads/main`), so they could not share a group in any case; | |
| # the condition states the intent where a reader of this file will look for it, rather than | |
| # leaving it as a property of GitHub's ref naming that a future group change could lose. | |
| concurrency: | |
| group: ${{ github.workflow }}-${{ github.ref }} | |
| cancel-in-progress: ${{ github.event_name == 'pull_request' }} | |
| # Workflow level: every job here runs tests or lints and reports, none writes to the | |
| # forge, and a job added later inherits read-only instead of the repository default of | |
| # read/write (#32). Three jobs as of #635 (pytest, shell, lint) -- this line names none | |
| # of them by count on purpose, so a fourth job does not make it stale again. | |
| permissions: | |
| contents: read | |
| jobs: | |
| pytest: | |
| runs-on: ${{ matrix.os }} | |
| # #1658: 15 left no headroom on windows-latest, where the suite itself finishes | |
| # (green) in ~862-873s (0:14:22-0:14:33) and checkout/setup/dependency-install | |
| # overhead pushes the job's total wall time past the ceiling -- three | |
| # back-to-back cancellations on the same green commit (jobs #105401634166, | |
| # #105406249723, #105410692355). 20 restores several minutes of real margin over | |
| # the slowest observed leg without removing the cap on a genuine hang. | |
| # | |
| # #1660: 20 was itself cancelled again, twice, on PR #1654's own commit | |
| # `fad72a25` (jobs #105421966394, #105427076121) -- the suite's own SELF- | |
| # REPORTED runtime (its own "N passed... in Xs" line) had grown from ~873s to | |
| # ~1176.55s (0:19:36), leaving only ~24s of job-wall-clock margin against the | |
| # 20-minute cap. Two things this file does NOT claim, having been caught | |
| # overclaiming both in review: whether that ~300s of growth is itself test | |
| # count/content growth, runner variance, or a genuine slowdown is unresolved | |
| # (open question, not "not more tests running"); and how long the process was | |
| # actually still alive after its own summary line printed is UNKNOWN, not | |
| # "several minutes" -- both runs were cut off by the cap itself, so the true | |
| # size of any post-summary gap was never observed, only that it is non-zero | |
| # (in both runs a `KeyboardInterrupt` lands in the main thread, blocked in a | |
| # `threading` wait/join, after the summary already reported every test | |
| # passing). Root cause not identified; #1660's own issue text asks for exactly | |
| # the diagnostic `tests/posthang_diagnostics_1660.py` now adds: periodic | |
| # thread-stack dumps once a session has finished but the process has not | |
| # exited. 30 is deliberately generous rather than a repeat of the same narrow | |
| # margin #1658 used, since the runtime has already grown once, unexplained, | |
| # and may again before the diagnostic pins down why. | |
| # | |
| # It did grow again, and the diagnostic did not help: a THIRD occurrence | |
| # (job #105442284149, commit `c201a8a0`, carrying #1660's own faulthandler | |
| # diagnostic) still ended `cancelled`, `1773.10s (0:29:33)` against this | |
| # 30-minute cap -- only ~26.9s of margin, and no `faulthandler` output | |
| # anywhere in that job's log. That is not the diagnostic failing to catch a | |
| # hang; it never got the chance to -- `POST_SESSION_DUMP_AFTER_SECONDS` was | |
| # 60 at the time, so the first dump was scheduled for 33s AFTER this cap | |
| # would already have killed the job. Fixed in | |
| # `tests/posthang_diagnostics_1660.py` (lowered to 15, with a coupling test | |
| # in `tests/test_posthang_diagnostics_1660.py` that catches THIS delay | |
| # going stale against `tests/test_pytest_leg_timeout_1658.py`'s own | |
| # hand-maintained runtime constant -- not a live measurement, so a FOURTH | |
| # occurrence still needs that constant updated by hand before the coupling | |
| # test can say anything about it). This cap is left at 30 rather than | |
| # raised a third time: the growth (873s -> 1176.55s -> 1773.10s) has tracked | |
| # each cap raise closely enough that raising it again is not expected to | |
| # hold either, and #1660's own issue text argues against repeating that fix | |
| # for however long the trend continues (this was true when this cap was | |
| # last left at 30; #1784 below raises it a third time anyway, but for a | |
| # different, now-measured reason -- see that comment). The next | |
| # recurrence, if the diagnostic can now actually fire, is what settles | |
| # whether this is a real hang or something else growing the suite's own | |
| # reported runtime. | |
| # | |
| # #1784: this 30-minute figure was clearing OBSERVED_WORST_CASE_SUITE_MINUTES | |
| # (`tests/test_pytest_leg_timeout_1658.py`) by only ~26.9s, and that whole | |
| # margin was computed against pytest's own self-reported runtime alone -- | |
| # `timeout-minutes` counts from JOB start, so checkout/setup-python/the | |
| # Windows Defender exclusion step/`pip install`, none of which pytest's own | |
| # clock ever sees, were silently eating into it. Measured directly from the | |
| # GitHub Actions API for the six jobs already cited above and in | |
| # `tests/posthang_diagnostics_1660.py`, that pre-pytest overhead ran 26-43s | |
| # -- already comparable to, and at its worst exceeding, the entire 26.9s | |
| # margin the arithmetic assumed was free. Raised by 2 minutes specifically | |
| # to cover that now-measured, largely fixed overhead (not to chase the | |
| # suite's own separate, unexplained runtime growth the paragraph above | |
| # already argues against re-raising for) -- see | |
| # `tests/test_pytest_leg_timeout_1658.py`'s own `PRE_PYTEST_OVERHEAD_SECONDS` | |
| # for the six measured values this raise is sized against. | |
| timeout-minutes: 32 | |
| # #1246: the `os` x `python-version` product below is the full 3x4 release matrix, | |
| # unconditionally -- so `full_matrix: true` restores full coverage by excluding | |
| # nothing, rather than needing a second, separately-maintained matrix to drift out | |
| # of sync with this one. An ordinary push/pull_request, or a workflow_dispatch left | |
| # at its `false` default, drops to the five legs #1246 measured and kept: ubuntu | |
| # 3.12 (cheap default), ubuntu 3.9 (the declared floor), windows 3.12 (the observed | |
| # failure platform), macos 3.12 (the third OS), plus the unconditional `shell` job | |
| # below. `exclude:` lists the eight combinations that reduction drops; each entry | |
| # must name both `os` and `python-version` or GitHub Actions treats it as a | |
| # wildcard on the field left out, excluding every python-version for that os (or | |
| # vice versa) instead of the one combination intended. | |
| strategy: | |
| fail-fast: false | |
| matrix: | |
| os: [ubuntu-latest, macos-latest, windows-latest] | |
| python-version: ["3.9", "3.10", "3.11", "3.12"] | |
| exclude: ${{ (github.event_name == 'workflow_dispatch' && github.event.inputs.full_matrix == 'true') && fromJson('[]') || fromJson('[{"os":"ubuntu-latest","python-version":"3.10"},{"os":"ubuntu-latest","python-version":"3.11"},{"os":"macos-latest","python-version":"3.9"},{"os":"macos-latest","python-version":"3.10"},{"os":"macos-latest","python-version":"3.11"},{"os":"windows-latest","python-version":"3.9"},{"os":"windows-latest","python-version":"3.10"},{"os":"windows-latest","python-version":"3.11"}]') }} | |
| env: | |
| # The exact CPython patch `setup-python` resolves "3.9" to on the hosted | |
| # windows-latest runner today. Both the cache path and the cache key below are | |
| # derived from it, and the disclosure step checks the interpreter that actually | |
| # ran against it -- two literals that must agree is a bug waiting for the day | |
| # one of them is edited without the other. | |
| WIN_PY39_PATCH: "3.9.13" | |
| steps: | |
| - uses: actions/checkout@v7 | |
| # CPython 3.9 has left the windows-latest image, so `setup-python` reports | |
| # `Version 3.9 was not found in the local cache` and downloads and installs it -- | |
| # measured at 44s of a 229s critical-path leg under xdist (#1196; #1177 landed | |
| # first and made this leg the critical path in the first place). Restoring the | |
| # hosted tool cache here is the whole fix: `setup-python` then finds the tree | |
| # and skips the download. | |
| # | |
| # `Digital-Process-Tools/claude-supertool` already ships this pattern for its | |
| # own CI (#1127 there); this step and its trap comment are ported from that | |
| # workflow (confirmed via `gh api | |
| # repos/Digital-Process-Tools/claude-supertool/contents/.github/workflows/tests.yml` | |
| # rather than retyped from a description). | |
| # | |
| # The cached path is the VERSION directory, not `.../x64` beneath it. | |
| # `@actions/tool-cache` records an installed tool by writing a sibling marker | |
| # `x64.complete` next to `x64/`, and its `find()` returns a miss when that | |
| # marker is absent -- so caching only `x64/` restores a tree `setup-python` | |
| # then ignores, and every run is a cache HIT that saves nothing. | |
| # | |
| # Windows-only and 3.9-only on purpose: the other three Windows legs already | |
| # have the interpreter baked into the image, and macOS pays a similar ~22s for | |
| # 3.9 but is never the critical path (107s against Windows' 266-305s), so | |
| # caching there would buy billed minutes and no wall clock. | |
| # | |
| # A cache written on a pull request branch is visible only to that branch: the | |
| # `push: [main]` run is what populates the key every other branch restores | |
| # from, so the first PR after this change lands may still miss -- expected, not | |
| # a failure. The "Tool cache state" step below discloses which happened rather | |
| # than leaving it to be inferred from `Set up Python`'s own duration. | |
| - name: Restore the CPython ${{ env.WIN_PY39_PATCH }} tool cache (#1196) | |
| id: toolcache-39 | |
| if: runner.os == 'Windows' && matrix.python-version == '3.9' | |
| uses: actions/cache@v6 | |
| with: | |
| path: ${{ runner.tool_cache }}/Python/${{ env.WIN_PY39_PATCH }} | |
| key: ${{ runner.os }}-${{ runner.arch }}-toolcache-python-${{ env.WIN_PY39_PATCH }} | |
| - name: Set up Python ${{ matrix.python-version }} | |
| uses: actions/setup-python@v7 | |
| with: | |
| python-version: ${{ matrix.python-version }} | |
| # Windows Defender's real-time scanner walks every file this job touches on the | |
| # runner's own disk, and the suite creates many small temp files per assertion | |
| # (each `mktemp` fixture, each rewritten .claude/jit-context/ under test). This is | |
| # a documented GitHub Actions cost specific to windows-latest: macOS and Linux | |
| # runners carry no equivalent always-on scanner. Excluding the checkout and the | |
| # runner temp directory removes that tax without touching what any test asserts -- | |
| # it changes nothing about correctness, only how fast the disk answers. Ported | |
| # from Digital-Process-Tools/claude-jit-context#310 (open, not yet merged there, | |
| # at the time this landed -- the shape below is that PR's actual diff, not a | |
| # retyped guess). The measured saving on this repo's own suite is pending a live | |
| # CI round; see #938. | |
| # | |
| # Each Add-MpPreference call is wrapped in try/catch rather than given a bare | |
| # continue-on-error (#1081): the cmdlet has been observed to fail on a hosted | |
| # runner with 0x800106ba (the Defender service itself unavailable), a cause | |
| # unrelated to the change under test, and an unguarded step failure there fails | |
| # the whole leg and renders identically to a real Windows test failure in | |
| # gh-pr:N:status -- distinguishable only by opening the job log. | |
| # continue-on-error alone would stop that but stays silent about which | |
| # happened; try/catch keeps the step from failing the job AND lets one place | |
| # disclose "applied" vs "failed" per path, reusing the hit/miss/skip | |
| # disclosure shape #1196 already put in this workflow (see "Tool cache state" | |
| # below) rather than inventing a second one. | |
| - name: Exclude the checkout and temp dir from Windows Defender scanning | |
| if: runner.os == 'Windows' | |
| shell: pwsh | |
| run: | | |
| foreach ($path in @("${{ github.workspace }}", "$env:RUNNER_TEMP")) { | |
| try { | |
| Add-MpPreference -ExclusionPath $path -ErrorAction Stop | |
| Write-Host "defender-exclusion: ok ($path excluded)" | |
| } catch { | |
| Write-Host "::warning::defender-exclusion: failed ($path not excluded -- $($_.Exception.Message))" | |
| } | |
| } | |
| # Three states, never a silent one. `cache-hit: true` on its own is not | |
| # evidence the cache did anything: `WIN_PY39_PATCH` is a literal, and if | |
| # `setup-python` ever resolves "3.9" to a different patch, the step above would | |
| # restore -- and then re-save -- a directory nothing reads. A permanent HIT | |
| # sitting next to a permanent 44s download is this repo's own defect class | |
| # exactly (an absence produced by the tool, read as an absence in the world), | |
| # so the check is against the interpreter that actually ran, not against the | |
| # cache action's own opinion of itself. | |
| # | |
| # `::error::` annotates the run summary without failing the step: a | |
| # runner-image change is not the fault of whoever pushes next. | |
| # | |
| # NOT `always()`: GitHub runs `bash` with `set -eo pipefail`, so if the step | |
| # above failed there is no guarantee `python` behaves the same way, and a step | |
| # that cannot run is not itself a finding -- it is reported as `skipped`. | |
| - name: Tool cache state (#1196) | |
| if: runner.os == 'Windows' | |
| shell: bash | |
| env: | |
| PY_VERSION: ${{ matrix.python-version }} | |
| CACHE_HIT: ${{ steps.toolcache-39.outputs.cache-hit }} | |
| run: | | |
| if [ "$PY_VERSION" != "3.9" ]; then | |
| echo "toolcache : skipped (CPython $PY_VERSION is baked into the windows-latest image; only 3.9 is missing from it)" | |
| exit 0 | |
| fi | |
| resolved=$(python -c "import sys; print('.'.join(map(str, sys.version_info[:3])))" || true) | |
| if [ -z "$resolved" ]; then | |
| echo "toolcache : skipped (could not read the interpreter version -- this says nothing about the cache either way)" | |
| exit 0 | |
| fi | |
| if [ "$resolved" != "$WIN_PY39_PATCH" ]; then | |
| echo "::error::toolcache stale pin - this workflow caches CPython $WIN_PY39_PATCH but setup-python resolved $resolved, so the restore can never be read. Set WIN_PY39_PATCH to $resolved in .github/workflows/tests.yml." | |
| echo "toolcache : finding (stale pin: cached $WIN_PY39_PATCH, running $resolved -- the restore is dead weight)" | |
| exit 0 | |
| fi | |
| if [ "$CACHE_HIT" = "true" ]; then | |
| echo "toolcache : ok (hit -- CPython $resolved restored, no download)" | |
| else | |
| echo "toolcache : miss (CPython $resolved downloaded; the tool cache is saved at job end, for the next run -- expected on the first run after this change on a given branch, since a cache written on a PR branch is visible only to that branch)" | |
| fi | |
| # markdown-it-py is not a test tool, it is the assembler's one dependency, and the | |
| # suite over scripts/assemble_changelog.py cannot reach a single assertion without | |
| # it -- the script refuses to run rather than fall back to text scanning. Absent | |
| # here, that whole suite errored on every leg, so the heading refusal, the raw-HTML | |
| # refusal and the fence state machine had never been exercised on a runner at all; | |
| # they passed locally because the package happened to be installed there. The | |
| # changelog workflow learned this the same way and the lesson was applied to that | |
| # one file only. tests/test_workflow_dependencies.py now holds both against | |
| # scaffold.ASSEMBLER_DEPENDENCIES, which is where the package name is declared. | |
| # pyyaml is here for tests/test_shell_leg_budget_303.py, which asserts against the | |
| # parsed `shell` job rather than against substrings of this file -- a regex over a | |
| # workflow is the shape that keeps passing while the workflow is broken. It is the | |
| # only test that needs it; the rest of the suite reads workflows as text on | |
| # purpose. That file fails rather than skips when `CI` is set and pyyaml is | |
| # missing, so dropping this package cannot make it go quiet. | |
| # | |
| # ruff is here (#635) so tests/test_ruff_ratchet_635.py actually runs instead of | |
| # skipping: that file's own `pytestmark` declines every test when `ruff` is not on | |
| # PATH, and this is the only job that runs `pytest` at all -- the `lint` job below | |
| # never does. Absent here, the ratchet script's own logic (baseline comparison, | |
| # exit-code selection, the third `could-not-run` state) would ship a whole new file | |
| # of tests that pass by skipping on all 12 legs of this matrix, every time. | |
| - name: Install test deps | |
| run: | | |
| python -m pip install --upgrade pip | |
| pip install pytest pytest-cov pytest-xdist markdown-it-py pyyaml ruff==0.16.3 | |
| # #1177: what `pytest -n auto` would request on THIS leg, from doctor.py's own | |
| # transcription of xdist's worker-count logic (#367), before any decision is taken | |
| # about parallelising this matrix. The runners' core counts have never been read | |
| # here -- the 2-3x figure in #1177 is reasoned from a documentation page, and | |
| # GitHub has changed those numbers before (public repositories went from 2 to 4 | |
| # vCPU in 2024). Twelve legs printing their own answer is the reading that | |
| # argument needs, and it is worth keeping afterwards: the number a later | |
| # regression would move is then in every run's log rather than in one issue. | |
| # | |
| # doctor.py exits 0 in every mode, so this can never fail a leg, and it reports | |
| # `worker sizing unknown` rather than a number when nothing on the runner | |
| # answered. Unconditional on purpose: a leg that skips it reports nothing, and | |
| # the platforms this matrix exists for are exactly the ones whose sizing is not | |
| # known here. psutil is deliberately not installed for it -- xdist would not find | |
| # one either, so the source this prints is the source `-n auto` would really use. | |
| - name: Report what `pytest -n auto` would request here | |
| run: python scripts/doctor.py --worker-sizing | |
| # #1177 candidate 1: `-n auto --dist loadfile`. The probe added by #1189 has now | |
| # reported on all three platforms of this matrix -- ubuntu 4 workers, windows 4, | |
| # macOS 3, ubuntu's read from three legs of run 34056088247 -- so the parallelism | |
| # available here is measured rather than taken from a documentation page. | |
| # | |
| # `--dist loadfile` and not the default `--dist load`: this suite's tests build git | |
| # repositories, change working directory and manage worktrees, and keeping one | |
| # file's tests on one worker is what lets those fixtures survive being run beside | |
| # each other. It is the weaker parallelism and it is the one worth trying first. | |
| # | |
| # What this does NOT buy, so nobody reads the worker count as a speedup factor: a | |
| # whole file lands on one worker, so the longest single FILE sets the floor. | |
| # `tests/test_no_test_pins_the_current_version_350.py` alone was 25.33s of ubuntu | |
| # 3.12's 234.23s of measured test time, and no number of workers divides that. | |
| # | |
| # The risk here is not wall clock, it is a suite that passes in a different order. | |
| # A green first run is the outcome to distrust rather than the one to ship: a race | |
| # that fires one run in ten arrives later as a flake nobody attaches to this change. | |
| # | |
| # What xdist did NOT cause, recorded because this change was blamed for it for two | |
| # rounds: the 123 launcher-suite failures on the Windows legs came from the | |
| # `shell: bash` the coverage step briefly carried, and reproduced on a serial run | |
| # with no workers in the log at all. See that step's own comment. | |
| # | |
| # #1176: 126 of the 428 tracked test files counted at filing time are content | |
| # guards over this repo's own markdown/config -- their answer does not vary by OS | |
| # or interpreter, and running | |
| # them on all 12 legs buys nothing a reader ever sees. Marking every one of them | |
| # was weighed and declined: the naive "no tmp_path/subprocess/monkeypatch" heuristic | |
| # selects the CHEAP tests (9.67s total, ran locally), not the expensive ones -- the | |
| # nine tests actually worth deselecting spawn a subprocess or read this repo's whole | |
| # tree, and #1176 marks exactly those nine `invariant` (tests/test_invariant_ | |
| # marker_1176.py holds the registry). `-m "not invariant"` on every leg but the full | |
| # one deselects them; the full leg is the SAME leg #1177 already designated to keep | |
| # coverage, `(ubuntu-latest, 3.12)`, so this reuses one special-leg concept instead | |
| # of inventing a second, and it keeps the deselect off the 51.6%-of-wall-clock | |
| # Windows legs. | |
| # | |
| # Both flags live on the same unconditional step, each behind its own `${{ }}` | |
| # ternary with the identical condition -- never two `if:`-guarded steps, because | |
| # two conditions that are both false is a leg that runs no tests and reports | |
| # success anyway (this repository's own defect class, installed in the merge gate | |
| # on purpose). Each ternary is written `X && 'flag' || ''` rather than | |
| # `X && '' || 'flag'`: the empty string is falsy in a GitHub expression, so the | |
| # non-empty arm has to be the TRUE arm or the flag lands on every leg instead of | |
| # eleven of them -- #1177's own fix for the identical mistake on the coverage flag. | |
| - name: Run tests | |
| run: | | |
| # #1673: any single test past 180s dumps its own stack and name to | |
| # stderr (pytest's built-in faulthandler plugin), so a hung test is | |
| # named in the log rather than showing as a bare KeyboardInterrupt | |
| # at the job's own cap -- which four issues read as a post-session | |
| # hang before this named the real one on its first run. | |
| pytest -n auto --dist loadfile -o faulthandler_timeout=180 ${{ (matrix.os != 'ubuntu-latest' || matrix.python-version != '3.12') && '--no-cov' || '' }} ${{ (matrix.os != 'ubuntu-latest' || matrix.python-version != '3.12') && '-m "not invariant"' || '' }} | |
| shell: | |
| runs-on: ubuntu-latest | |
| timeout-minutes: 10 | |
| steps: | |
| - uses: actions/checkout@v7 | |
| # Both guards below scoped themselves with `git ls-files '*.sh'`, which returns | |
| # exactly one path here -- scripts/doctor.sh. bin/oss-workspace is tracked, is | |
| # POSIX sh and has no extension, so neither `bash -n` nor `shellcheck` had ever | |
| # read the plugin's own entry point, on any leg, on any platform, in any release, | |
| # and this leg was green throughout: a lint that ran and found nothing and a lint | |
| # that never received the file both exit 0 (#193). | |
| # | |
| # The list is derived by extension OR shebang, in scripts/shell_sources.py, and | |
| # the reasoning for deriving it rather than naming the file here is in that | |
| # script's docstring. Two properties this job depends on: | |
| # | |
| # * it exits 2 when it matches nothing, so an empty selection fails this leg | |
| # instead of running `shellcheck` with no arguments and exiting 0 -- the half | |
| # of the repair that would otherwise reintroduce the bug; | |
| # * it exits 3 when a tracked file could not be read, because `could not look` | |
| # is not `looked and found no shebang`. | |
| # | |
| # Redirection and not a pipe: the default shell here is `bash -e` without | |
| # `pipefail`, so `... | tee list` would report tee's exit status and swallow both | |
| # refusals. | |
| - name: Enumerate shell sources | |
| run: | | |
| python3 scripts/shell_sources.py --root . > shell-sources.txt | |
| echo "linting:"; cat shell-sources.txt | |
| # `while ... done < file` and not `... | while`: a pipeline puts the loop in a | |
| # subshell, and a `while` returns the status of the loop's completion rather than | |
| # of its body, so a failing guard would be swallowed. The status is therefore | |
| # carried out by hand -- `|| exit 1` below, and a flag in the shellcheck step so | |
| # every file is reported rather than only the ones before the first failure. | |
| # | |
| # `xargs -d` would be shorter and is a GNU extension; this form is the one that | |
| # can be run and watched fail on a maintainer's own machine, which is how the | |
| # `| while` trap above was caught rather than reasoned about. | |
| - name: Syntax-check every shell source | |
| run: | | |
| while IFS= read -r f; do bash -n "$f" || exit 1; done < shell-sources.txt | |
| # No package fetch here, and that is the whole of #303 rather than a tidy-up. | |
| # | |
| # This step used to open with `sudo apt-get update -qq && sudo apt-get install -y | |
| # shellcheck`. Measured on job 96152482222 (run 32276977038, main at 7ea64c9) the | |
| # job ran 10.32 minutes against `timeout-minutes: 10` and was killed. GitHub | |
| # renders a timeout kill as `cancelled`, not `failure`, so main and every open | |
| # pull request read as "0 failed, 1 cancelled -- not green" with nothing broken, | |
| # six times in one day. | |
| # | |
| # Two readings decided this rather than a bigger cap: | |
| # | |
| # * the killed step logged nothing at all between its own `##[endgroup]` and | |
| # `##[error]The operation was canceled.` 10m12s later. `apt-get install` is | |
| # not quiet, so it was never reached: the entire window was `apt-get update`; | |
| # * every successful run logged `shellcheck is already the newest version | |
| # (0.9.0-1)`. The install has never installed anything -- shellcheck ships in | |
| # the ubuntu-latest image -- while `apt-get update` cost 68.4s of a 70s step | |
| # on 2026-08-19 and 5.2s on 2026-08-16. The linting itself is under a second. | |
| # | |
| # So the fetch is removed rather than pinned. A pinned release tarball or a cached | |
| # apt layer would each be a network round trip bought for a binary already on the | |
| # runner, and pinning one means carrying a version and a checksum for it. | |
| # | |
| # `timeout-minutes` is deliberately left at 10 and is now a bound on a hang rather | |
| # than a budget this job routinely spends: the measured job is a few seconds. It | |
| # is not raised, because the number it was failing against was almost entirely | |
| # `apt-get`; it is not lowered, because nothing measured argues for a particular | |
| # smaller number and a tight cap reintroduces reds caused by nothing in a diff. | |
| # | |
| # The version now floats with the runner image. That is the cost, and it is made | |
| # visible rather than assumed: the version is printed, and the binary being absent | |
| # is its own exit code. Before this, an absent shellcheck exited 127 per file into | |
| # the same `fail` flag a real finding uses, so `could not lint` and `linted and | |
| # found a problem` reached the leg's status identically -- and `A && B` under | |
| # `bash -e` does not exit on A, so a failed install fell through to exactly that. | |
| - name: shellcheck | |
| run: | | |
| if ! command -v shellcheck > /dev/null 2>&1; then | |
| echo "shellcheck is not on PATH." >&2 | |
| echo "It ships in the ubuntu-latest image and this job deliberately fetches" >&2 | |
| echo "nothing (#303). If the image has stopped carrying it, install it here" >&2 | |
| echo "again -- but pin a version and a checksum, and put the fetch in a step" >&2 | |
| echo "of its own so a slow mirror is not charged to this lint's budget." >&2 | |
| exit 4 | |
| fi | |
| shellcheck --version | |
| fail=0 | |
| while IFS= read -r f; do shellcheck -S warning "$f" || fail=1; done < shell-sources.txt | |
| exit "$fail" | |
| # #635 -- 250 tracked .py files had no linter anywhere. The write-time half | |
| # (supertool's `ruff` validator, `.supertool.json`, catches a finding the | |
| # moment a file is edited) cannot see a file nobody touches in this run, and | |
| # is not wired into CI at all, so this leg is what covers the whole tree and | |
| # cannot be skipped. Ubuntu-only and its own job (not folded into `pytest` | |
| # above): the ruleset in pyproject.toml's `[tool.ruff.lint]` does not vary | |
| # by OS or interpreter, so running it 12 times would be 11 wasted runs | |
| # buying no new coverage -- the same reasoning `shell` above already uses. | |
| # | |
| # `scripts/ruff_ratchet.py` is a ratchet, not a "must be zero" gate: #635's | |
| # own measurement found 95 pre-existing findings across 51 files outside | |
| # that issue's claimed scope, so this only fails a PR that makes the count | |
| # go UP, never one that leaves it flat. See that script's own docstring. | |
| lint: | |
| runs-on: ubuntu-latest | |
| timeout-minutes: 10 | |
| steps: | |
| - uses: actions/checkout@v7 | |
| - name: Set up Python | |
| uses: actions/setup-python@v7 | |
| with: | |
| python-version: "3.12" | |
| - name: Install ruff | |
| run: | | |
| python -m pip install --upgrade pip | |
| pip install ruff==0.16.3 | |
| - name: ruff (#635 ratchet) | |
| run: python3 scripts/ruff_ratchet.py --root . |