Skip to content

PR Readiness

PR Readiness #2943948

Workflow file for this run

name: PR Readiness
# One current-SHA verdict for the otherwise wide PR workflow fan-out. The
# commit status is the branch-protection handle; the label is the at-a-glance
# signal on PR lists and the PR header.
on:
# This workflow never checks out or executes PR code. pull_request_target
# gives the trusted base workflow enough permission to replace a stale label
# immediately for both same-repository and fork PRs.
pull_request_target:
types:
- opened
- synchronize
- reopened
- edited
- ready_for_review
- converted_to_draft
- closed
workflow_run:
workflows:
- CI
# The cheap blocking gates in fast-gate.yml, split out of CI so the fork reviewers can
# key on them. It is a lane in its own right -- a red gate must red the PR --
# and its completion is also what releases CI's heavy jobs.
- Fast Gate
- Build
- Code Review
# Blocking: a PR that declares no triaged issue reds it. Runs on
# `pull_request` for every PR, forks included, so its payload carries
# the PR head like Code Review's does.
- Issue Gate
- Opus 5.5 Review
- GPT 6.1 Review
- Design Review
- UX Review
- First Principles Review
# Blocking and fail-closed: a script-confirmed newly-refused operation reds
# it, and so does a run that measured nothing.
- Security Scope Review
- CodeQL
# A workflow that is itself `workflow_run`-triggered CANNOT be listed here.
# GitHub runs such a workflow from the default branch, so the
# `workflow_run` payload its completion hands this job carries the default
# branch as head_branch and the default branch's tip as head_sha -- never
# the pull request's head. `Resolve current pull request revision` then
# asks which open pull request has `<this repo>:<default branch>` as its
# head, an answer that is empty by construction, and exits SKIP without
# publishing anything. The stage-2 fork reviewers (`Fork * Review`,
# `Fork Internal Content Scan`) all key on Fast Gate this way, so each of
# their completions dispatches a readiness run that can only no-op: 96% of
# this workflow's run creation and one wasted REST request each against the
# shared hourly pool (see docs/ci/ci-and-reviews.md).
#
# A fork pull request's verdict is refreshed the same way every other
# verdict is: `pr-readiness-sweep.yml` re-fires the recompute by PR
# number when a check on its head completes, and is therefore immune to
# this payload problem.
# `test_ai_review_workflows.py` pins the rule so a future lane cannot be
# added back here without the same silent no-op.
#
# It runs on `pull_request`, so its payload does carry the PR head and
# its start can hold the verdict like any other lane's.
- Internal Content Scan Gate
# ONLY `in_progress`. `completed` is deliberately absent, and so is
# `requested`.
#
# A `workflow_run` dispatch is spent the moment GitHub delivers the event,
# before this job runs a step: a run that then finds nothing to do has
# already cost a runner start and its queue slot. Every type listed here
# fires once per monitored workflow per revision, so with `completed`
# listed this workflow created 24 runs per head update (12 x 2) plus one
# per re-run attempt -- measured at 30 per head and 3000 an hour, the
# majority of this repository's run creation, on a job whose real work
# takes 9-16 seconds. The per-revision concurrency group below collapses
# that burst for EXECUTION, but a collapsed run has already consumed its
# dispatch, so the group never bounded this cost. Neither can anything a
# step does: the only lever on dispatch is which events are subscribed.
#
# So the terminal verdict is delivered by `pr-readiness-sweep.yml` instead.
# It scans every open pull request over GraphQL -- a SEPARATE pool from the
# REST budget the lanes share -- and dispatches this workflow, by PR number
# and head SHA, for exactly the heads on which a monitored check completed
# after the current verdict was published. One scan replaces the
# per-completion fan-out: 37 heads carried monitored completions in a
# measured 15-minute window, against 282 completion events. The cost is
# latency, bounded by the sweep's cadence plus its publish lag, and only
# in the SAFE direction: a verdict goes green later than the event would
# have made it, never earlier.
#
# Why `in_progress` stays, and why it is the one type that cannot move to
# the sweep: it is the only signal that a monitored workflow went BACK to
# running. A re-run reuses the same workflow run and increments its
# attempt, so `requested` never fires for it and the runs page still shows
# the pre-re-run conclusion until the attempt finishes. Without it,
# readiness holds the pre-re-run `success` for the whole re-run -- and
# since this status is the branch-protection handle for the entire
# fan-out, armed auto-merge can merge a revision whose lane is at that
# moment failing. A sweep cannot close that: its evidence is COMPLETED
# checks, and a running lane has none yet. The event does, in seconds.
#
# What an `in_progress` run does is therefore deliberately minimal: it
# resolves the pull request, confirms the run the event names is still
# running (a hold that arrives after the lane completed would stamp the
# status newer than that completion and hide it from the sweep), and if
# so publishes `pending` and stops (see "Hold the verdict" below). No
# evaluation, and no read of the status it is about to replace -- a
# read-then-decide on the STATUS would reintroduce the stale-success
# window the publish step's own comments document -- so the whole run is
# three requests. The lane's completion then reaches the sweep as
# evidence, and the recompute it dispatches publishes the real verdict.
#
# Why `requested` carries nothing: it fires when a run is CREATED, which
# on a head update happens for all 12 workflows above before any of them
# can have a verdict -- and readiness has already published `checking` for
# that revision via the `pull_request_target: synchronize` path.
#
# The per-head count above only holds for a monitored workflow whose
# payload names the pull request's own head. A `workflow_run`-triggered
# lane names the default branch instead, so its dispatches key on ONE
# shared group (the default branch's tip) and accumulate across every open
# pull request rather than per head update -- measured at 158 readiness
# runs on a single default-branch SHA, none of which can resolve a pull
# request. Such a lane is barred from the allowlist above for that reason.
# Measurements and the queue-saturation mechanism: docs/ci/ci-and-reviews.md.
types: [in_progress]
workflow_dispatch:
inputs:
pr:
description: "Pull request number to refresh"
required: true
type: string
sha:
description: "Expected current pull request head SHA"
required: true
type: string
permissions:
actions: read
checks: read
contents: read
issues: write
pull-requests: write
statuses: write
concurrency:
# pull_request_target runs are the only readiness runs that surface as a
# CheckRun in the PR's status rollup. If they share a cancelling group -- with
# the workflow_run burst OR with each other -- a superseded PT run is marked
# "cancelled" by GitHub, and that cancelled check then shows on the PR even
# though the authoritative "PR Readiness" commit status is fine (this is the
# "canceled at the end of CI" symptom: a workflow_run fires as a lane starts
# and collapses an in-flight PT run). GitHub always marks a
# superseded run cancelled -- cancel-in-progress:true kills the in-flight one,
# and false still cancels the pending one -- so the only way a PT run never
# shows cancelled is to keep it out of any group that can cancel it. Each PT
# run therefore gets its own isolated group and always resolves to a real
# conclusion (success / failure / skipped), never cancelled. The evaluate and
# publish steps are idempotent and stale-SHA guarded, so un-collapsed PT runs
# on superseded revisions simply no-op green. workflow_run / workflow_dispatch
# runs do NOT appear in the rollup, so they keep the cheap per-revision burst
# collapse via cancel-in-progress. A workflow_run burst collapses on
# (head repository, head branch, head SHA) and NOT on the PR number, because
# `workflow_run.pull_requests` is empty whenever the head repository is a
# fork -- keying on it degraded every fork burst to one group per run. That
# triple identifies the same PR, since only one open PR can exist per source
# repository + branch.
#
# An `in_progress` run still cancels a running evaluation, and that is the
# merge guard's other half: the cancelled run's reads predate the re-run, so
# it must not publish. The hold this run writes instead stands until the
# sweep's recompute, which reads the re-run's completion.
group: >-
${{
github.event_name == 'pull_request_target'
&& format('pr-readiness-pt-{0}', github.run_id)
|| github.event_name == 'workflow_dispatch'
&& format('pr-readiness-wd-{0}-{1}', inputs.pr, inputs.sha)
|| format('pr-readiness-wr-{0}-{1}-{2}',
github.event.workflow_run.head_repository.full_name
|| github.run_id,
github.event.workflow_run.head_branch || '',
github.event.workflow_run.head_sha
|| github.run_id)
}}
cancel-in-progress: true
jobs:
readiness:
name: Publish readiness signal
# This gate MUST NOT consult `workflow_run.pull_requests`: that array is
# empty whenever the head repository is a fork, so requiring a number in it
# skipped every fork PR's re-evaluation. A fork PR's verdict was then only
# ever computed by the pull_request_target run -- which fires BEFORE the
# monitored workflows exist -- and the commit status stayed pending forever
# with no later event able to recompute it. The head SHA is the fork-safe
# key, and the step below resolves it to a PR (skipping cleanly when no
# open PR owns it), so admitting every pull_request run is correct and the
# non-PR events are already excluded by the event check.
#
# A `workflow_run.event == 'workflow_run'` upstream is NOT admitted. Such a
# run was dispatched from the default branch, so its payload names the
# default branch's tip as the head and can never resolve to a pull request
# (see the allowlist note above). Admitting it only reached the resolve
# step's empty lookup. The allowlist is the primary fence -- no monitored
# workflow is `workflow_run`-triggered any more -- and this is the second:
# re-adding such a lane has to clear both, which
# `test/test_ai_review_workflows.py` refuses.
if: >-
github.event_name != 'workflow_run'
|| github.event.workflow_run.event == 'pull_request'
|| github.event.workflow_run.event == 'dynamic'
runs-on: ubuntu-latest
timeout-minutes: 20
env:
# The same-repo workflow-run lanes of the verdict step's spec list, read
# by the publish step's pre-success re-check.
MONITORED_LANES: >-
ci.yml fast-gate.yml build.yml code-review.yml issue-gate.yml
internal-content-scan-gate.yml claude-review.yml codex-review.yml
design-review.yml ux-review.yml first-principles-review.yml
security-scope-review.yml
# Old check-run name -> current name for the fork lanes #16238 renamed,
# read by the publish step's fork re-check (the one read that binds a row
# by name). Why and how: docs/ci/ci-and-reviews.md. Remove once no open
# PR has a pre-#16238 head (#16373).
LEGACY_LANE_NAMES: >-
{"Opus 5 Review": "Opus 5.5 Review", "GPT 5.6 Review": "GPT 6.1 Review"}
steps:
# Transient network/TLS errors between the runner and api.github.com can
# abort a single-shot `gh` call, and under `set -e` that would end the
# whole evaluation before the publish step could run, leaving a red
# check unrelated to actual readiness. Every read-only gh call whose
# failure would end the evaluation goes through this bounded helper: 3
# attempts with short backoff, mirroring the bounded-attempts convention
# in publish-installer.yml. Two reads deliberately do not: the existing
# PR Readiness status read uses `bounded` alone and falls back to
# `__unreadable__`, and the rate-limit log line is best-effort. Stdout
# is buffered per attempt so a failed attempt's partial output (e.g. an aborted --paginate stream) never
# leaks into a pipe. WRITE calls are deliberately NOT retried here; the
# label writes have their own narrow race tolerance (see drop_label /
# ensure_label), and a lost-response retry of a DELETE outside that
# tolerance would turn an applied change into a new failure.
- name: Install bounded retry helper for read-only gh calls
run: |
cat > "$RUNNER_TEMP/gh-retry.sh" <<'HELPER'
# `timeout` is GNU coreutils: present on the runner, ABSENT on macOS by
# default -- and these step scripts are extracted and executed verbatim
# by test_pr_readiness_evaluate.py / _publish.py, so they have to run
# there too. Missing, `timeout` makes every call exit 127, `gh_retry`
# reads that as a transient failure, and the lane publishes `pending`
# for a reason that has nothing to do with the PR under test.
#
# Resolved once here rather than at each call site, so the 120s bound
# also has a single definition. Falling back to UNBOUNDED is deliberate:
# a call with no ceiling is a weaker guarantee than a bounded one, but a
# far better answer than treating the tool's absence as the request
# having failed. Written for bash 3.2, which is what macOS ships.
bounded() {
if command -v timeout >/dev/null 2>&1; then
timeout 120 "$@"
elif command -v gtimeout >/dev/null 2>&1; then
gtimeout 120 "$@"
else
"$@"
fi
}
gh_retry() {
local attempt out err rc=1
for attempt in 1 2 3; do
# `bounded` caps a HANGING attempt (gh applies no HTTP request
# timeout of its own); rc 124 is just another retryable failure.
if out="$(bounded "$@" 2>"$RUNNER_TEMP/gh-retry-err")"; then
cat "$RUNNER_TEMP/gh-retry-err" >&2
printf '%s\n' "$out"
return 0
else
rc=$?
fi
err="$(cat "$RUNNER_TEMP/gh-retry-err" 2>/dev/null || true)"
printf '%s\n' "$err" >&2
# A non-429 HTTP 4xx is a PERMANENT request error (renamed
# workflow file, missing scope, bad endpoint): retrying cannot
# fix it, and calling it "transient" would misdiagnose a
# misconfiguration as network trouble. Return a distinct code
# immediately so callers can fail loud instead of publishing a
# pending "could not be evaluated" verdict forever.
# EXCEPTION: GitHub's primary/secondary rate limits surface as
# HTTP 403 with rate-limit text in the body -- those are
# transient exactly like a 429 and must be retried, not failed.
if grep -Eq 'HTTP 4[0-9][0-9]' <<<"$err" \
&& ! grep -q 'HTTP 429' <<<"$err" \
&& ! grep -Eiq 'rate limit|abuse detection' <<<"$err"; then
echo "gh_retry: permanent HTTP error (no retry): $*" >&2
return 22
fi
# The PRIMARY limit is the exception to the exception. It is
# the hourly budget every workflow in this repository shares
# through GITHUB_TOKEN, and it refills at the top of the hour,
# not in the seconds this loop waits -- so attempts 2 and 3 are
# certain to fail and only add to the request volume that
# emptied the pool. GitHub words it "API rate limit exceeded
# for <installation|user|...>"; the secondary limit says
# "secondary rate limit" and IS worth a short backoff. Return
# the transient code without retrying: the evaluate step turns
# it into the non-terminal "could not be evaluated" verdict
# and the next event (or the sweep) recomputes once the hour
# has turned.
if grep -q 'API rate limit exceeded' <<<"$err"; then
echo "gh_retry: primary rate limit exhausted (no retry): $*" >&2
return "$rc"
fi
echo "gh_retry: attempt $attempt/3 failed (rc=$rc): $*" >&2
if [ "$attempt" -lt 3 ]; then
sleep $(( attempt * 2 ))
fi
done
return "$rc"
}
HELPER
- name: Resolve current pull request revision
id: context
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
REPO: ${{ github.repository }}
EVENT: ${{ github.event_name }}
INPUT_PR: ${{ inputs.pr }}
INPUT_SHA: ${{ inputs.sha }}
EVENT_PR: ${{ github.event.pull_request.number }}
EVENT_SHA: ${{ github.event.pull_request.head.sha }}
RUN_PR: ${{ github.event.workflow_run.pull_requests[0].number }}
RUN_SHA: ${{ github.event.workflow_run.head_sha }}
RUN_HEAD_REPO: ${{ github.event.workflow_run.head_repository.full_name }}
RUN_HEAD_REF: ${{ github.event.workflow_run.head_branch }}
EVENT_DEFAULT_BRANCH: ${{ github.event.repository.default_branch }}
run: |
set -euo pipefail
source "$RUNNER_TEMP/gh-retry.sh"
case "$EVENT" in
workflow_dispatch) PR="$INPUT_PR"; EXPECTED_SHA="$INPUT_SHA" ;;
workflow_run)
PR="$RUN_PR"
EXPECTED_SHA="$RUN_SHA"
# `workflow_run.pull_requests` is empty for a fork head AND for
# the managed `dynamic` CodeQL runs, so it can never be the only
# way in. Resolve the PR from the immutable head SHA instead.
#
# NOT via `repos/{repo}/commits/{sha}/pulls`: that endpoint only
# lists PRs whose head commit is reachable in THIS repo, and a
# fork PR's head commit is not reachable from any base-repo
# branch until it merges -- so it returns empty for every OPEN
# fork PR, skipping every re-evaluation and freezing the fork
# verdict at the `opened` snapshot (a premature red). Match the
# head SHA against the open-PR list instead: a PR object always
# carries its `head.sha` regardless of which repo the commit
# lives in, so this resolves fork and same-repo heads alike.
#
# A head SHA is not by itself unique -- two open PRs can point at
# the same commit (e.g. the same fork commit pushed to two
# branches). Disambiguate with the triggering run's head
# repository + branch (both populated on every workflow_run),
# exactly as the verdict step binds monitored runs below.
#
# Ask for THAT branch, not for every open PR. `head=<owner>:<ref>`
# is the list endpoint's own filter, so the answer is one page
# of (almost always) one PR instead of a walk over the whole
# open set -- at 600 open PRs that walk was seven requests on
# the pool every lane shares, on every fork trigger. The owner
# is the head repository's owner: a fork PR's head lives under
# the fork owner, and a same-repo PR's under this repository's
# owner. The filter cannot express the repository NAME, so a
# user with two forks under one account would match both; the
# `.head.repo.full_name` check below keeps the right one. The
# SHA-only walk stays as the fallback for the event shape that
# carries neither field.
if [ -z "$PR" ]; then
select_pr='
first(.[][]
| select(.head.sha == $sha
and ($repo == "" or .head.repo.full_name == $repo)
and ($ref == "" or .head.ref == $ref))
| .number) // empty
'
if [ -n "$RUN_HEAD_REPO" ] && [ -n "$RUN_HEAD_REF" ]; then
head_q="$(jq -rn --arg v "${RUN_HEAD_REPO%%/*}:$RUN_HEAD_REF" '$v | @uri')"
PR="$(gh_retry gh api \
"repos/$REPO/pulls?state=open&head=$head_q&per_page=100" \
| jq -rs \
--arg sha "$RUN_SHA" \
--arg repo "$RUN_HEAD_REPO" \
--arg ref "$RUN_HEAD_REF" "$select_pr")"
else
PR="$(gh_retry gh api --paginate \
"repos/$REPO/pulls?state=open&per_page=100" \
| jq -rs \
--arg sha "$RUN_SHA" \
--arg repo "$RUN_HEAD_REPO" \
--arg ref "$RUN_HEAD_REF" "$select_pr")"
fi
fi
if [ -z "$PR" ]; then
{
echo "state=SKIP"
echo "stale=true"
} >> "$GITHUB_OUTPUT"
echo "No open pull request currently has workflow head $RUN_SHA."
exit 0
fi
;;
pull_request_target) PR="$EVENT_PR"; EXPECTED_SHA="$EVENT_SHA" ;;
*) echo "::error::unsupported event: $EVENT"; exit 1 ;;
esac
pr_json="$(gh_retry gh pr view "$PR" --repo "$REPO" \
--json number,state,isDraft,isCrossRepository,baseRefName,headRefName,headRefOid,headRepository,headRepositoryOwner,url)"
# The PR's OWN author, for the conditions that exempt a bot's pull
# request. `github.actor` cannot serve them here: the terminal verdict
# is delivered by the sweep's `workflow_dispatch`, whose actor is
# `github-actions[bot]` however the PR was opened, so an actor test
# exempts nothing on exactly the run that decides.
#
# Read over REST rather than taking `gh pr view --json author`: every
# dependabot condition in this repository compares against
# `dependabot[bot]`, which is REST's `.user.login`. `gh pr view` goes
# through GraphQL, where a Bot actor's login is the bare name
# (`dependabot`) -- add-contributor.yml's `__typename == "Bot"` note is
# the same distinction -- so normalising one spelling into the other
# would be an inference. One extra read is cheaper than a comparison
# that silently stops matching.
PR_AUTHOR="$(gh_retry gh api "repos/$REPO/pulls/$PR" --jq '.user.login')"
SHA="$(jq -r '.headRefOid' <<<"$pr_json")"
STATE="$(jq -r '.state' <<<"$pr_json")"
# The default branch is where CI / Build / CodeQL are gated to run
# (`pull_request: branches: [<default>]` and CodeQL default-setup).
# A PR whose base is any other branch -- a stacked PR sitting on
# another feature branch -- can never start those lanes, so the
# verdict step marks them ineligible instead of pending-forever.
# Every trigger here carries it in `github.event.repository`, so the
# request is only the fallback for a payload that does not.
DEFAULT_BRANCH="${EVENT_DEFAULT_BRANCH:-}"
if [ -z "$DEFAULT_BRANCH" ]; then
DEFAULT_BRANCH="$(gh_retry gh api "repos/$REPO" --jq '.default_branch')"
fi
{
echo "pr=$PR"
echo "sha=$SHA"
echo "state=$STATE"
echo "fork=$(jq -r '.isCrossRepository' <<<"$pr_json")"
echo "draft=$(jq -r '.isDraft' <<<"$pr_json")"
echo "url=$(jq -r '.url' <<<"$pr_json")"
echo "base=$(jq -r '.baseRefName' <<<"$pr_json")"
echo "author=$PR_AUTHOR"
echo "default_branch=$DEFAULT_BRANCH"
# (head repository, head branch) is how a monitored workflow run is
# bound back to this PR below. Both fields are populated on a fork
# run, unlike `pull_requests`.
echo "head_repo=$(jq -r '
"\(.headRepositoryOwner.login)/\(.headRepository.name)"
' <<<"$pr_json")"
echo "head_ref=$(jq -r '.headRefName' <<<"$pr_json")"
} >> "$GITHUB_OUTPUT"
if [ -n "$EXPECTED_SHA" ] && [ "$EXPECTED_SHA" != "$SHA" ]; then
echo "stale=true" >> "$GITHUB_OUTPUT"
echo "Ignoring stale $EVENT event for $EXPECTED_SHA; PR #$PR is now at $SHA."
else
echo "stale=false" >> "$GITHUB_OUTPUT"
fi
- name: Remove readiness labels from closed pull request
if: steps.context.outputs.stale != 'true' && steps.context.outputs.state != 'OPEN'
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
REPO: ${{ github.repository }}
PR: ${{ steps.context.outputs.pr }}
run: |
set -euo pipefail
source "$RUNNER_TEMP/gh-retry.sh"
# A peer run of this workflow can remove the same label between our
# snapshot and our DELETE, and a label that is already gone has
# reached the state we wanted. Treat only that 404 as success; any
# other failure is still a failure.
drop_label() {
local label="$1" encoded out
encoded="$(jq -rn --arg value "$label" '$value | @uri')"
if out="$(gh api --method DELETE \
"repos/$REPO/issues/$PR/labels/$encoded" 2>&1)"; then
return 0
fi
if grep -q 'HTTP 404' <<<"$out"; then
echo "Label '$label' was already removed by a concurrent run."
return 0
fi
printf '%s\n' "$out" >&2
return 1
}
current="$(gh_retry gh api "repos/$REPO/issues/$PR/labels" --paginate --jq '.[].name')"
for label in \
"readiness: checking" \
"readiness: action required" \
"readiness: maintainer review" \
"readiness: passed"; do
if grep -Fxq "$label" <<<"$current"; then
drop_label "$label"
fi
done
# The disposition gate below runs a repository script, so the bytes it
# executes must come from TRUSTED code -- this workflow is
# pull_request_target and holds write tokens, so checking out the PR head
# would hand a contributor arbitrary execution with them. Pin the DEFAULT
# BRANCH explicitly rather than relying on the per-event default ref
# (base branch for pull_request_target, default branch for workflow_run):
# one visible ref is what makes the trust boundary reviewable.
# github.event.repository is absent on some payloads, so fall back to
# github.ref_name, which is the default branch on those triggers.
# sparse-checkout keeps this to the one script directory, and
# persist-credentials: false leaves no token in the worktree -- nothing
# here runs git against the remote.
- name: Hold the verdict while a monitored lane runs
id: hold
if: >-
github.event_name == 'workflow_run'
&& github.event.action == 'in_progress'
&& steps.context.outputs.stale != 'true'
&& steps.context.outputs.state == 'OPEN'
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
REPO: ${{ github.repository }}
SHA: ${{ steps.context.outputs.sha }}
URL: ${{ steps.context.outputs.url }}
LANE: ${{ github.event.workflow_run.name }}
RUN_ID: ${{ github.event.workflow_run.id }}
run: |
set -euo pipefail
source "$RUNNER_TEMP/gh-retry.sh"
READ_FAILURE_TOKEN="[read-failed]"
# The event said a monitored lane went back to running. This run may
# execute long after it was dispatched -- queue delay is the very
# condition this workflow's dispatch count feeds -- so first ask
# whether that is STILL true, of the run the event names. A lane
# that has since completed needs no hold: its completion is the
# sweep's evidence, and a hold written now would stamp the status
# NEWER than that completion, pushing it below the sweep's evidence
# floor with no later event left to lift it. That is the freeze the
# sweep exists to end, manufactured by the guard meant to prevent
# it, on exactly the saturated days the guard matters most.
#
# This read is of the workflow run, not of the commit status: it
# asks whether the fact being acted on still holds, never what the
# status currently says. Reading the status first is the shape the
# publish step's POST comment records three guards failing to make
# safe, and it is not done here.
hold_description="readiness check(s) changed since evaluation; re-checking ${LANE:-a monitored lane}"
if lane_status="$(gh_retry gh api "repos/$REPO/actions/runs/$RUN_ID" --jq '.status')"; then
if [ "$lane_status" = "completed" ]; then
echo "pr-readiness: ${LANE:-the lane} already completed on $SHA before this hold ran; writing nothing, the sweep reads its completion"
exit 0
fi
else
rc=$?
if [ "$rc" -eq 22 ]; then exit 1; fi
# Unknown whether the lane is still running, so the hold is written
# -- pending only ever blocks -- and stamped as a read failure. The
# sweep retries that class on age alone, so a hold that did land
# over the lane's completion is recomputed regardless of evidence.
hold_description="$READ_FAILURE_TOKEN could not confirm ${LANE:-a monitored lane} is still running; holding $SHA until re-evaluated"
fi
# `pending`, written unconditionally once the lane is known (or not
# known) to be running: a commit status is last-write-wins per
# SHA+context with no conditional write, so nothing is gained by
# reading it first. A pending over a pending only moves the stamp; a
# pending over a failure blocks the merge just the same, and the lane
# that produced the failure is by definition being re-run.
#
# This run stops here. The lane's completion reaches the sweep as
# check evidence newer than this stamp, and the recompute it
# dispatches publishes the real verdict with the labels to match.
# Labels are advisory decoration and are left to that recompute:
# the status is the branch-protection handle, and it is written
# FIRST and alone for the same reason the publish step writes it so.
jq -n \
--arg target_url "$URL" \
--arg description "$hold_description" \
'{
state: "pending",
target_url: $target_url,
description: $description,
context: "PR Readiness"
}' > "$RUNNER_TEMP/readiness-hold.json"
if ! bounded gh api --method POST "repos/$REPO/statuses/$SHA" \
--input "$RUNNER_TEMP/readiness-hold.json" >/dev/null; then
echo "::error::Failed to hold the readiness status for $SHA while" \
"${LANE:-a monitored lane} re-runs. Not retried (a retry can" \
"overwrite a concurrent run's newer verdict); the sweep" \
"recomputes when the lane completes."
exit 1
fi
echo "held=true" >> "$GITHUB_OUTPUT"
echo "pr-readiness: holding $SHA at pending while ${LANE:-a monitored lane} runs"
- name: Check out the disposition evaluator from the default branch
if: >-
steps.context.outputs.stale != 'true'
&& steps.context.outputs.state == 'OPEN'
&& steps.hold.outputs.held != 'true'
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
ref: ${{ github.event.repository.default_branch || github.ref_name }}
sparse-checkout: src/kiro_crew/builtin_skills/kirocrew-dev/kirocrew-prepare-pr/scripts
persist-credentials: false
# Issue #6658: the one-lane / one-rationale-per-finding disposition rule
# was mechanical only for a writer running the kirocrew-prepare-pr loop, so a
# writer who skips that loop could post a blanket single-rationale
# `target=gpt` record that codex-review.yml's adjudication ledger admits
# with full downgrade power while nothing on the merge path objected.
# Evaluating it here binds every writer, because this job publishes the
# repository's only required status.
#
# The evaluation is NOT reimplemented in shell: `--disposition-gate` runs
# the same `disposition_violations` the local gate calls, over the same
# records the ledger admits, so the rule keeps ONE definition instead of
# gaining a workflow-side fourth copy of the grammar.
#
# This step must never fail the job. A failed step skips the publish step
# below, which leaves the required status pending with no further event to
# recompute it on an unchanged commit -- the permanent-pending failure
# mode this workflow is written around. So: no `set -e`, every outcome
# becomes an output, and an unreadable record set becomes `ok=false`,
# which the evaluation reads as WAITING, never as a red.
#
# Cost, stated plainly: this is the repository's highest-dispatch
# workflow, and the gate adds a shallow sparse checkout plus one comment
# read (and one permission call per distinct disposition author, which on
# a PR with no disposition comments is zero). It cannot be skipped on a
# cheap precondition, because deciding whether a disposition comment
# exists means reading the comment list and matching the marker prefix --
# a fourth reading of the grammar, which is the thing #6658 exists to
# avoid. The per-revision concurrency group already collapses the
# workflow_run burst, so the added work lands roughly once per head
# update per uncollapsed event, not once per monitored workflow.
- name: Evaluate disposition records
id: dispositions
if: >-
steps.context.outputs.stale != 'true'
&& steps.context.outputs.state == 'OPEN'
&& steps.hold.outputs.held != 'true'
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
REPO: ${{ github.repository }}
PR: ${{ steps.context.outputs.pr }}
SHA: ${{ steps.context.outputs.sha }}
GATE: ${{ github.workspace }}/src/kiro_crew/builtin_skills/kirocrew-dev/kirocrew-prepare-pr/scripts/pr_status.py
run: |
set -uo pipefail
ok=false
violations=""
unpublished=""
if [ ! -f "$GATE" ]; then
# The evaluator moved or the sparse checkout missed it. Report
# UNKNOWN so readiness waits and says why, rather than reporting a
# clean rule from a check that never ran.
echo "pr-readiness: disposition evaluator not found at $GATE" >&2
elif out="$(python3 "$GATE" --disposition-gate \
--repo "$REPO" --pr "$PR" --head "$SHA")"; then
echo "pr-readiness: disposition gate -- $out"
if [ "$(jq -r '.ok' <<<"$out" 2>/dev/null)" = "true" ]; then
ok=true
violations="$(jq -r '.violations[]' <<<"$out" 2>/dev/null || true)"
unpublished="$(jq -r '.unpublished[]' <<<"$out" 2>/dev/null || true)"
else
echo "pr-readiness: disposition gate could not evaluate: $(jq -r '.error' <<<"$out" 2>/dev/null)" >&2
fi
else
# stderr above is deliberately NOT suppressed: the gate turns every
# expected failure into ok=false JSON, so reaching here means the
# interpreter itself failed and the traceback is the only diagnosis.
echo "pr-readiness: disposition gate did not run (exit $?)" >&2
fi
# Random heredoc delimiter per GitHub's own guidance for multi-line
# outputs: the violation text is charset-limited (logins, span ids,
# lane names) and newline-flattened by the gate, so this is belt and
# braces rather than the only guard.
delim="EOF_$(head -c 16 /dev/urandom | od -An -tx1 | tr -d ' \n')"
{
echo "ok=$ok"
echo "violations<<$delim"
printf '%s\n' "$violations"
echo "$delim"
echo "unpublished<<$delim"
printf '%s\n' "$unpublished"
echo "$delim"
} >> "$GITHUB_OUTPUT"
# Issue #13951: a review lane's marker comment is ONE slot selected by the
# lane marker alone, never by the head, so a second sample at one head
# REPLACES the first and only the survivor is presented. The direction that
# matters is a blocking sample replaced by a clean one at the same head: the
# `[BLOCK-MERGE]` line goes, the stamp still names the head, and this status
# goes green over a block nobody adjudicated away.
#
# Evaluated HERE for the same reason the disposition rule is: a reader that
# only runs in the local kirocrew-prepare-pr loop binds whoever runs that loop, and
# #13951's own finding was that the record existed in GraphQL
# userContentEdits while nothing read it. A reader no gate calls reproduces
# that shape one layer up -- the reading exists and nothing consumes it.
#
# SCOPE, deliberately the narrowest thing that closes the hole: only
# `blocking_dropped` gates. It fires only when a replaced sample carried
# `[BLOCK-MERGE] <head>` AND the body now presented is itself stamped for
# that head and does not block. Measured over 559 bot comments with 26
# same-head replacement pairs: ONE true detection and no false positive, so
# gating on it does not redden an ordinary revision. (An earlier scan of the
# same condition reported zero occurrences; it was blind to three of the
# five lanes, because it read only the `[BLOCK-MERGE]` marker and the
# whole-design lanes write `<Lane>-Verdict: BLOCK` instead. That zero is
# retracted -- do not reason from it.) `superseded` alone is deliberately
# NOT a gate: ordinary same-head re-samples happen, are not harmful, and
# reddening them would block real revisions for a condition this repository
# does not treat as a defect.
#
# Never fails the step, same as the disposition gate: a failed step skips
# the publish below and strands the required status pending with no event to
# recompute it.
#
# Cost, stated plainly: one comment read plus one GraphQL edit-history read
# per bound lane that posted (at most five). Like the disposition gate it
# cannot be skipped on a cheap precondition -- deciding whether any lane
# comment was ever replaced means reading the comment list and then its
# history, which is the work itself.
# Skipped for `dependabot[bot]` for the same reason every review lane skips
# it, and with the same condition those lanes use: on such a pull request no
# bound lane ever publishes a marker comment, so this gate examines nothing
# and correctly reports UNKNOWN. Mapping that UNKNOWN to `pending` would
# strand the required status for the life of the head -- and the sweep cannot
# rescue it, because its rescue needs a check that completed after the
# verdict was published, while here every recompute re-derives the same empty
# reading and there is no lane to re-run that would ever post a comment. The
# exemption is deliberately the ACTOR and not the empty reading: a head whose
# lanes have simply not posted YET is also empty, and that one must stay
# UNKNOWN, because its lanes will publish and can then be superseded.
# Leaving the step unrun makes its outputs empty, which the evaluation below
# treats as contributing nothing -- the same shape the disposition step uses.
- name: Evaluate superseded verdicts
id: supersessions
if: >-
steps.context.outputs.stale != 'true'
&& steps.context.outputs.state == 'OPEN'
&& steps.hold.outputs.held != 'true'
&& steps.context.outputs.author != 'dependabot[bot]'
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
REPO: ${{ github.repository }}
PR: ${{ steps.context.outputs.pr }}
SHA: ${{ steps.context.outputs.sha }}
GATE: ${{ github.workspace }}/src/kiro_crew/builtin_skills/kirocrew-dev/kirocrew-prepare-pr/scripts/pr_status.py
run: |
set -uo pipefail
ok=false
dropped=""
if [ ! -f "$GATE" ]; then
echo "pr-readiness: supersession evaluator not found at $GATE" >&2
elif out="$(python3 "$GATE" --supersession-gate \
--repo "$REPO" --pr "$PR" --head "$SHA")"; then
echo "pr-readiness: supersession gate -- $out"
if [ "$(jq -r '.ok' <<<"$out" 2>/dev/null)" = "true" ]; then
ok=true
dropped="$(jq -r '.blocking_dropped[]' <<<"$out" 2>/dev/null || true)"
else
echo "pr-readiness: supersession gate could not evaluate: $(jq -r '.error' <<<"$out" 2>/dev/null)" >&2
fi
else
# stderr deliberately not suppressed: the gate turns every expected
# failure into ok=false JSON, so reaching here means the interpreter
# itself failed and the traceback is the only diagnosis.
echo "pr-readiness: supersession gate did not run (exit $?)" >&2
fi
delim="EOF_$(head -c 16 /dev/urandom | od -An -tx1 | tr -d ' \n')"
{
echo "ok=$ok"
echo "dropped<<$delim"
printf '%s\n' "$dropped"
echo "$delim"
} >> "$GITHUB_OUTPUT"
- name: Evaluate current revision
id: verdict
if: >-
steps.context.outputs.stale != 'true'
&& steps.context.outputs.state == 'OPEN'
&& steps.hold.outputs.held != 'true'
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
REPO: ${{ github.repository }}
PR: ${{ steps.context.outputs.pr }}
SHA: ${{ steps.context.outputs.sha }}
HEAD_REPO: ${{ steps.context.outputs.head_repo }}
HEAD_REF: ${{ steps.context.outputs.head_ref }}
DRAFT: ${{ steps.context.outputs.draft }}
FORK: ${{ steps.context.outputs.fork }}
BASE_REF: ${{ steps.context.outputs.base }}
DEFAULT_BRANCH: ${{ steps.context.outputs.default_branch }}
TRIGGER_EVENT: ${{ github.event_name }}
TRIGGER_ACTION: ${{ github.event.action }}
# The triggering workflow_run's own name + status, read ONLY to close
# a race on the fork Stage-2 lanes below: this evaluation can itself be
# the `in_progress` event a human-override rerun fires the INSTANT it
# starts, before that rerun's own "Open check-run" step has posted a
# fresh row -- so a check-runs read right now would still see the OLD
# completed verdict and publish a stale success. Empty on every other
# trigger (pull_request_target, or a workflow_run this isn't).
WR_NAME: ${{ github.event.workflow_run.name }}
WR_STATUS: ${{ github.event.workflow_run.status }}
# From the disposition step above. Empty DISPOSITION_OK means the step
# did not run at all, which contributes nothing -- the step's own `if`
# matches this one's, and test_pr_readiness_dispositions.py pins the
# wiring so a rename cannot silently disable the gate.
DISPOSITION_OK: ${{ steps.dispositions.outputs.ok }}
DISPOSITION_VIOLATIONS: ${{ steps.dispositions.outputs.violations }}
# The whole-design lanes the same gate call found owing this head a
# verdict they never published -- see the advisory-slot note below.
ADVISORY_UNPUBLISHED: ${{ steps.dispositions.outputs.unpublished }}
# From the supersession step above, on the same terms: empty means the
# step did not run, which contributes nothing. Only the dropped-block
# list gates; test_prepare_pr_supersession.py pins these two names to
# this step's outputs so a rename cannot silently empty the env and
# leave both `case` arms falling through with every test green.
SUPERSESSION_OK: ${{ steps.supersessions.outputs.ok }}
SUPERSESSION_DROPPED: ${{ steps.supersessions.outputs.dropped }}
run: |
set -euo pipefail
source "$RUNNER_TEMP/gh-retry.sh"
# Set when a read-only API call keeps failing after the bounded
# retries: the verdict is then UNKNOWN, not red -- see the explicit
# non-terminal branch after the evaluation loop.
transport_failed=false
# CI, Build and CodeQL only ever run for a PR whose base is the
# default branch: ci.yml / build.yml declare `pull_request:
# branches: [<default>]`, and CodeQL default-setup only scans PRs
# into the default branch. A stacked PR (base is another feature
# branch) therefore never starts those workflows, so the head SHA
# carries no run for them. Counting a run that structurally cannot
# exist as "(not started)" froze such a PR at pending forever, with
# no future event able to recompute it. Mark them ineligible on a
# non-default base instead -- the same "Not eligible" treatment the
# fork lane gives CodeQL. This is not a coverage hole: when the
# bottom PR of the stack merges, GitHub auto-retargets this PR to the
# default branch and CI/Build/CodeQL then run over the combined diff
# before anything can reach the default branch.
stacked=false
if [ -n "$DEFAULT_BRANCH" ] && [ "$BASE_REF" != "$DEFAULT_BRANCH" ]; then
stacked=true
fi
workflow_specs=()
skipped=()
if [ "$stacked" = "true" ]; then
skipped+=("CI (only runs on PRs to $DEFAULT_BRANCH)")
skipped+=("Fast Gate (only runs on PRs to $DEFAULT_BRANCH)")
skipped+=("Build (only runs on PRs to $DEFAULT_BRANCH)")
else
workflow_specs+=(
"ci.yml|CI"
# Fast Gate carries the same `branches:` filter as CI, so it belongs
# in this carve-out for the same reason: on a stacked PR it never
# starts, and a lane that reads "(not started)" freezes the verdict
# at pending forever.
"fast-gate.yml|Fast Gate"
"build.yml|Build"
)
fi
# Code Review and the AI reviewers trigger on `pull_request` with no
# base-branch filter, so they run on every PR regardless of base.
workflow_specs+=("code-review.yml|Code Review")
# Issue Gate runs on `pull_request` with no base filter and needs
# only a read token, so it runs on every PR too, forks and stacked
# PRs included. It is enrolled HERE, not as a second branch-protection
# entry: `PR Readiness` is the one required status, so a lane this
# list omits is a gate that can go red without reddening the PR.
workflow_specs+=("issue-gate.yml|Issue Gate")
if [ "$FORK" = "true" ]; then
# A fork head cannot run the repository's managed default-setup
# CodeQL, so that lane stays ineligible. The AI code reviews DO run
# on forks: the privileged Stage-2 fork-*-review.yml pipeline fires
# on CI completion from the default branch and posts a check-run
# named exactly like the same-repo lane ("Opus 5.5 Review", etc.).
# Read those results from the head SHA's check-runs (checkrun:
# specs) so a fully-reviewed green fork can reach "passed" instead
# of being forced to a maintainer-review verdict.
skipped+=("CodeQL (fork PR)")
# The third field binds each verdict to THIS pull request AND this
# review attempt via the external_id the lane stamps: without it a
# sibling fork PR on the same head SHA, or a stale row from a
# previous rerun attempt, could answer for this one. The fourth
# field is the workflow whose completion triggers the lane -- all
# five fire on `Fast Gate` completing, so its newest run + attempt
# for this head is what "current" means here (fast-gate.yml).
workflow_specs+=(
# The scan reaches a fork through the same Stage-2 pattern: a fork
# head gets no OIDC token, so fork-internal-content-scan.yml runs
# privileged from the default branch and posts a check-run named
# exactly like the same-repo lane. Read it here so a fork PR is
# gated too -- external contributions are the LAST place to leave
# without pre-merge coverage. The third field binds the verdict to
# THIS pull request AND this scan attempt via the external_id the
# lane stamps: without it a sibling fork PR on the same head SHA,
# or a stale row from a previous rerun attempt, could answer for
# this one. The fourth field is the workflow whose completion
# triggers the lane -- its newest run + attempt is what "current"
# means here.
"checkrun:Internal Content Scan|Internal Content Scan|internal-content-scan-pr-|fast-gate.yml"
"checkrun:Opus 5.5 Review|Opus 5.5 Review|opus-pr-|fast-gate.yml"
"checkrun:GPT 6.1 Review|GPT 6.1 Review|gpt-pr-|fast-gate.yml"
"checkrun:Design Review|Design Review|design-pr-|fast-gate.yml"
"checkrun:UX Review|UX Review|ux-pr-|fast-gate.yml"
"checkrun:First Principles Review|First Principles Review|first-principles-pr-|fast-gate.yml"
"checkrun:Security Scope Review|Security Scope Review|scope-pr-|fast-gate.yml"
)
else
if [ "$stacked" = "true" ]; then
skipped+=("CodeQL (only runs on PRs to $DEFAULT_BRANCH)")
skipped+=("Internal Content Scan (only runs on PRs to $DEFAULT_BRANCH)")
else
workflow_specs+=("dynamic/github-code-scanning/codeql|CodeQL")
# Blocking lane: an added line carrying an Amazon-internal marker
# fails this and therefore fails readiness. Resolved by workflow
# FILE, so the awkward reusable-workflow check name
# ("scan / internal-content-scan") is not restated here.
workflow_specs+=("internal-content-scan-gate.yml|Internal Content Scan")
fi
workflow_specs+=(
"claude-review.yml|Opus 5.5 Review"
"codex-review.yml|GPT 6.1 Review"
"design-review.yml|Design Review"
"ux-review.yml|UX Review"
"first-principles-review.yml|First Principles Review"
"security-scope-review.yml|Security Scope Review"
)
fi
# ONE stable token marks EVERY pending this job publishes because a
# READ failed, rather than because a lane is genuinely still owed. The
# two are opposites for the self-heal sweep: a lane-pending is
# reproducible, so a recompute re-derives the same verdict, while a
# read-failure pending is not, so a recompute whose reads succeed
# resolves it outright. Only the description crosses to the sweep, so
# the token rides there -- FIRST, because the API caps a description
# at 140 characters and cuts the tail.
#
# A token, not a sentence: pr-readiness-sweep.yml matches this literal,
# and a sentence match recognises only the one phrasing it is written
# against, so a read-failure pending worded differently is stranded at
# pending with nothing but a push able to clear it. Every read-failure
# pending sets read_failure=true, and test_pr_readiness_sweep.py fails
# a site that does not.
READ_FAILURE_TOKEN="[read-failed]"
read_failure=false
passed=()
pending=()
failed=()
# A fork run with conclusion `action_required` was created but never
# executed: only a maintainer can approve it. Keep that condition out
# of `failed[]` so the verdict attributes the block correctly, and
# out of `pending[]` because automation cannot clear it by itself.
awaiting_approval=()
# ONE read for every workflow-run lane. Every `pull_request` run on
# this head comes back in a single page (19 workflows in this
# repository trigger on pull_request, one run each per head; a rerun
# re-uses the run and bumps its attempt), so the per-lane reads
# below are jq selects over it rather than one request per lane.
# Read per lane, this step cost ~11 requests per evaluation at
# ~250 evaluations an hour under load -- the largest single draw on
# the hourly pool every workflow here shares through GITHUB_TOKEN.
# That pool ran dry on 2026-09-15 and again on 2026-09-23, with
# every AI lane failing closed and this step logging
# "API rate limit exceeded for installation". `--paginate` is the
# ceiling guard: a head that somehow carries more than a page of
# runs still reads completely.
#
# A monitored file that does not exist is caught in CI, not here:
# test_pr_readiness_evaluate.py pins every `.yml` spec to a file
# under .github/workflows/, where the per-lane endpoint's loud 404
# used to catch a rename at run time. The consolidated list cannot
# tell "renamed" from "not started yet", and a runtime check would
# cost the request this read exists to save.
#
# A transport failure here is the one-lane transport failure it
# replaces, observed before any lane: the loop runs over no specs
# and the non-terminal "could not be evaluated" verdict publishes,
# exactly as it does for a `break` on the first spec.
pr_runs='{"workflow_runs":[]}'
if [ "${#workflow_specs[@]}" -gt 0 ]; then
pr_runs="$(gh_retry gh api --method GET --paginate \
"repos/$REPO/actions/runs?event=pull_request&head_sha=$SHA&per_page=100" \
| jq -cs '{workflow_runs: [.[].workflow_runs[]]}')" || {
rc=$?
if [ "$rc" -eq 22 ]; then exit 1; fi
transport_failed=true
workflow_specs=()
}
fi
# The head SHA's check-runs, read once and lazily: only a fork
# evaluation has checkrun: specs, and all seven read the same
# collection. Paginated because a head carries one check-run per
# JOB (89 on a typical same-repo head), and the second page is the
# one that would otherwise drop the lane posted last.
sha_check_runs=""
# The three whole-design lanes' comment slots, read ONCE here beside
# the runs read: both readers below score those lanes, and the
# same-repo lane and its fork counterpart share one slot each, so six
# lanes are answered by three rows.
#
# Why the slot is read at all: a lane that computed a verdict it could
# not publish still completes `success`. Every publish-failure arm of
# the lanes' shared upsert emits an `::error::` and then returns 0, so
# the run reports success while the slot still holds the PREVIOUS
# head's verdict. Counting that lane as passed states the design was
# reviewed for a head that carries no published review, and reporting
# it `failed` would publish a BLOCK verdict no reviewer reached. It is
# scored as a third outcome instead: a named pending that a re-run of
# that lane clears.
#
# Bound stamp names, never any `*-REVIEWED` match. A stamp for another
# lane's name inside this lane's body is model output, so only the
# slot's OWN name answers for it.
#
# NONE of those rules is spelled here. `$GATE`'s
# `evaluate_reviewer_markers` decides, per bound lane, whether that
# lane's own slot carries a stamp naming this head: the name binding,
# the abbreviated and elided stamp forms, the slots left alone for
# carrying no stamp of their own (the advisory lanes rewrite theirs to
# a stampless "skipped" notice by design, and a lane that published
# nothing at all is answered by its own conclusion), and the
# human-override record that stands in for a stamp on the path where
# the model is deliberately not re-run. That function is what
# `pr_status.py` reports BLOCKED from, so taking the answer from it is
# what makes the two gates unable to disagree about a head -- one
# definition, the same reason the disposition rule above is not
# restated in shell.
#
# It costs no extra request: that gate call already reads this PR's
# trusted comments for the disposition rule.
advisory_unpublished="${ADVISORY_UNPUBLISHED:-}"
# 1 when this advisory lane owes this head a verdict and has none,
# which is the state a `success` conclusion cannot express. All this
# adds to the gate's answer is which reviewer name the board's lane
# label belongs to.
advisory_slot_unpublished() {
local name
case "$1" in
"Design Review") name="DESIGN" ;;
"UX Review") name="UX" ;;
"First Principles Review") name="FIRST-PRINCIPLES" ;;
*) printf '0'; return 0 ;;
esac
if printf '%s\n' "$advisory_unpublished" | grep -Fqx -- "$name"; then
printf '1'
else
printf '0'
fi
}
# `${arr[@]+"${arr[@]}"}` rather than `"${arr[@]}"`: on the transport
# failure path above the array is EMPTY, and bash < 4.4 (macOS ships
# 3.2) treats expanding an empty array under `set -u` as an unbound
# variable and aborts the script -- which turned the non-terminal
# "could not be evaluated" verdict into rc=1 wherever this script
# runs under the system bash (the macOS test lane extracts and runs
# it). The `+` form expands to nothing when the array is unset or
# empty and to every element otherwise, on every bash.
for spec in ${workflow_specs[@]+"${workflow_specs[@]}"}; do
# Four fields: <workflow file or checkrun:NAME>|<label>|<external_id
# prefix>|<triggering workflow file>. The last two are OPTIONAL and
# only a checkrun: spec uses them -- see the binding note in that
# branch. Split on `|` alone so a label keeps its spaces.
IFS='|' read -r file label extid trigfile <<<"$spec"
if [[ "$file" == checkrun:* ]]; then
# Fork AI-review lane: the verdict lives in the head SHA's
# check-runs, not a PR-bound workflow run (the fork reviewer runs
# via workflow_run from the default branch, so its workflow-run
# head ref does not bind back to the fork). The same-repo lane
# publishes a DIFFERENT name on a fork, so read this name's
# check-runs, bound below to THIS PR and attempt via external_id,
# then collapse them: any still running -> pending; any
# failure-class -> fail; else any success/neutral -> pass; only
# an absent/unbound row (the real review has not posted yet) ->
# pending.
cname="${file#checkrun:}"
# A human-override-rerun race this branch does NOT currently
# close, kept here so the intent is not lost with the wiring.
# It fires only when readiness was itself triggered by a
# `Fork <lane>` run's `workflow_run: in_progress`, and such a
# trigger can never reach this step: a `workflow_run`-triggered
# lane runs from the default branch, so `Resolve current pull
# request revision` finds no open pull request for that head and
# exits SKIP long before the evaluation (see the trigger
# allowlist). The branch was therefore unreachable while those
# lanes were listed, and stays unreachable now that they are not
# -- and doubly so since every `in_progress` trigger stops at the
# hold step and never evaluates. The race it names is real and
# unguarded: a rerun's fresh "Open check-run" step can land after
# this read, so the read sees the OLD completed verdict. The
# sweep's mode-2 check evidence catches it within one sweep
# interval. Tracked in issue #14166 rather than fixed here, which
# is a trigger change only.
if [ "$FORK" = "true" ] && [ "${WR_STATUS:-}" = "in_progress" ] \
&& [ "${WR_NAME:-}" = "Fork $cname" ]; then
pending+=("$label (rerun starting)")
continue
fi
if [ -z "$sha_check_runs" ]; then
sha_check_runs="$(gh_retry gh api --method GET --paginate \
"repos/$REPO/commits/$SHA/check-runs?per_page=100" \
| jq -cs '{check_runs: [.[].check_runs[]]}')" || {
rc=$?
# 22 = permanent HTTP error (misconfiguration): fail loud so it
# is diagnosed, never diluted into "transient network trouble".
if [ "$rc" -eq 22 ]; then exit 1; fi
transport_failed=true
break
}
fi
# No name filter on the read: the external_id match below
# already names the lane (each lane stamps its own prefix), and
# it is the only binding a fork verdict is trusted on anyway.
cruns="$sha_check_runs"
# Bind the verdict to THIS pull request AND THIS review attempt
# when the lane stamps an external_id.
#
# Two open PRs can share a head SHA -- the same fork branch opened
# against two bases, or the same commit in two forks -- and each
# lane posts a check-run under this SAME NAME on that SHA. And a
# RERUN on an unchanged head leaves the previous completed row in
# place while the new attempt is still starting up. Reading by name
# alone loses both ways: a sibling's clean row, or a stale clean
# row from the previous attempt, answers for a review that has not
# reported. A model verdict can differ between attempts on the same
# head, so a stale clean verdict is not merely old, it can be WRONG.
#
# The fourth spec field names the workflow whose completion triggers
# the lane (Fast Gate, for all five). Its newest run for this head,
# plus that run's attempt, is what "the current attempt" means -- so
# the expected id is <prefix><pr>-<run id>-<attempt> and nothing
# else matches. When the expected row is not there yet, total=0 ->
# pending, which holds the merge instead of borrowing an answer.
#
# A human-override rerun (`gh api .../runs/<id>/rerun`) re-executes
# the LANE's run directly, without Fast Gate re-running, so the
# trigger-bound id stays identical between the stale failed attempt
# and the fresh rerun -- both rows carry the SAME external_id. That
# is safe: the match below collapses every row sharing that id to
# the newest by CHECK-RUN id (distinct per POST, monotonically
# increasing, independent of what external_id it carries), so the
# fresh rerun is never outvoted by the stale attempt it replaces.
#
# Run id AND attempt: a rerun keeps the id and bumps the attempt, so
# the id alone would still accept the stale row.
if [ -n "${extid:-}" ]; then
want="$extid$PR"
if [ -n "${trigfile:-}" ]; then
# Bind the trigger run to THIS PR by (head repository, head
# branch) before picking the newest, for the same reason the
# workflow-run branch below does it: the query is pinned to the
# head SHA, but two open PRs can share that SHA, so an unfiltered
# `max_by(.id)` can select the SIBLING's trigger run. The
# expected external_id would then name a run this PR's lane never
# saw, nothing would ever match, and the lane would sit pending
# forever -- the exact opposite failure to the stale-verdict one
# the binding exists to prevent, and just as bad.
trun="$(jq -c --arg path ".github/workflows/$trigfile" \
--arg repo "$HEAD_REPO" --arg ref "$HEAD_REF" '
[.workflow_runs[]
| select(.path == $path
and (.head_repository.full_name // "") == $repo
and (.head_branch // "") == $ref)]
| max_by(.id) // empty
' <<<"$pr_runs")"
if [ -z "$trun" ]; then
# The trigger has not run for this head yet, so no review is
# owed and none can have reported. Pending, not passed.
pending+=("$label (not started)")
continue
fi
want="$want-$(jq -r '.id' <<<"$trun")-$(jq -r '.run_attempt // 1' <<<"$trun")"
fi
# Match the exact trigger-bound id, then collapse to the
# newest CHECK-RUN by id (distinct per POST, monotonically
# increasing -- independent of the external_id these rows
# share). A lane rerun triggered by a human override
# (`gh api .../runs/<id>/rerun`) re-executes the lane directly
# without Fast Gate re-running, so the trigger-bound id stays
# identical between the stale failed attempt and the fresh
# one; the check-run id collapse alone still picks the fresh
# row, because each POST -- stale or fresh -- got its own
# higher check-run id regardless of what external_id it
# carries.
cruns="$(jq -c --arg x "$want" \
'{check_runs: [
[.check_runs[] | select(.external_id == $x)]
| max_by(.id) // empty
] | map(select(. != null))}' \
<<<"$cruns")"
fi
total="$(jq '.check_runs | length' <<<"$cruns")"
incomplete="$(jq '[.check_runs[]
| select(.status != "completed")] | length' <<<"$cruns")"
fail="$(jq '[.check_runs[]
| select(.status == "completed"
and (.conclusion | IN("failure","timed_out","cancelled",
"action_required","stale",
"startup_failure")))] | length' <<<"$cruns")"
pass="$(jq '[.check_runs[]
| select(.status == "completed"
and (.conclusion | IN("success","neutral")))] | length' <<<"$cruns")"
if [ "$total" -eq 0 ] || [ "$incomplete" -gt 0 ]; then
pending+=("$label (not started)")
elif [ "$label" = "Design Review" ] || [ "$label" = "UX Review" ] || [ "$label" = "First Principles Review" ]; then
# Design, UX and First Principles all BLOCK on a real BLOCK
# verdict, so a contributor cannot merge past a design judged
# wrong, an experience judged broken/misleading, or a surface
# judged unjustified. Safe to treat a failure as a blocker
# because each fork-*-review.yml fails the check ONLY on BLOCK
# -- an errored, throttled or verdict-less run resolves NEUTRAL
# and exits 0 -- so `failure` here means a judged-wrong
# verdict, never infrastructure noise, with one deliberate
# exception: the fork UX lane completes as `failure` when an
# attachment download from the description failed for a
# transport reason, a `cannot evaluate` that a re-run of that
# lane clears. A false BLOCK clears with `/ai-review override`:
# the handler records the judgment and re-runs the bound
# Stage-2 lane run (located through the lane-run marker in its
# check-run's output), whose "Resolve human override" step
# reads the record for this exact head, skips the model, and
# completes this row `success` with the override note in the
# slot -- the same outcome as the same-repo lane. This branch
# therefore reads that row like any other.
if [ "$fail" -gt 0 ]; then
failed+=("$label (BLOCK)")
elif [ "$pass" -gt 0 ]; then
if [ "$(advisory_slot_unpublished "$label")" = "1" ]; then
pending+=("$label (verdict not published, re-run this lane)")
else
passed+=("$label")
fi
else
pending+=("$label (not started)")
fi
# Security Scope Review is deliberately NOT in the branch above.
# That branch exists for lanes that resolve NEUTRAL when the run
# itself errored, so a `failure` there can only be a judged-wrong
# verdict. This lane fails CLOSED instead: an unreadable report, a
# missing head marker and a stage that never ran all red it, and
# each of those means the tightening went unmeasured. Relabeling
# its failure as "(BLOCK)" would claim a verdict was reached, and
# moving it out of `failed[]` would let an unmeasured tightening
# merge. The default branch below already attributes it as a plain
# blocker, which is true of both of its failing causes.
elif [ "$fail" -gt 0 ]; then
failed+=("$label (failure)")
elif [ "$pass" -gt 0 ]; then
passed+=("$label")
else
pending+=("$label (not started)")
fi
continue
fi
if [ "$file" = "dynamic/github-code-scanning/codeql" ]; then
runs="$(gh_retry gh api --method GET \
"repos/$REPO/actions/runs?event=dynamic&head_sha=$SHA&per_page=100")" || {
rc=$?
if [ "$rc" -eq 22 ]; then exit 1; fi
transport_failed=true
break
}
# Collapse to one run -- newest by monotonic id. Full rationale
# on the pull_request branch below; the two collapses must stay
# identical.
run="$(jq -c --arg path "$file" '
[.workflow_runs[] | select(.path == $path)]
| max_by(.id) // empty
' <<<"$runs")"
src="$runs"
else
# Bind the run to this PR by (head repository, head branch), NOT
# by `.pull_requests`: that array is empty on every fork run, so
# filtering on it matched nothing and reported each already-green
# workflow as "(not started)" -- a permanent pending verdict. The
# query is already pinned to this exact head SHA, and only one
# open PR can exist per source repository + branch.
# Collapse to one run on the monotonic run id, never on
# created_at: that timestamp has one-second granularity, and two
# runs of the same workflow on the same head routinely share a
# second (e.g. a workflow that fires on both synchronize and
# edited). jq's sort_by is stable and the API returns runs
# newest-first, so a created_at sort left same-second ties in
# API order and `last` picked the OLDEST run -- when that run
# was concurrency-cancelled, a lane whose newest run succeeded
# read as a blocking failure. max_by(.id) orders same-second
# runs correctly because ids increase monotonically. A cancelled
# newest run yields only to a LATER-STARTED run (below), so
# there is deliberately no plain cancelled-run filter: dropping
# it would revive an older verdict -- a manually cancelled rerun
# must read failure-class, not resurface a stale green.
run="$(jq -c --arg path ".github/workflows/$file" \
--arg head_repo "$HEAD_REPO" --arg head_ref "$HEAD_REF" '
[.workflow_runs[]
| select(.path == $path
and .head_repository.full_name == $head_repo
and .head_branch == $head_ref)]
| max_by(.id) // empty
' <<<"$pr_runs")"
src="$pr_runs"
fi
# Fork approval can start runs out of id order, so the max-id twin gets
# cancelled; the latest-started same-head run wins (ties: whole seconds).
if [ "$(jq -r '.conclusion // ""' <<<"${run:-null}")" = "cancelled" ]; then
run="$(jq -c --argjson n "$run" '
[.workflow_runs[]
| select(.path == $n.path and .head_branch == $n.head_branch
and .head_repository.full_name == $n.head_repository.full_name
and .id != $n.id and .run_started_at != null
and .run_started_at >= $n.run_started_at)]
| max_by([.run_started_at, .conclusion != "cancelled"]) // $n
' <<<"$src")"
fi
if [ -z "$run" ]; then
pending+=("$label (not started)")
continue
fi
status="$(jq -r '.status // ""' <<<"$run")"
conclusion="$(jq -r '.conclusion // ""' <<<"$run")"
if [ "$status" != "completed" ]; then
pending+=("$label ($status)")
else
if [ "$file" = "dynamic/github-code-scanning/codeql" ]; then
case "$conclusion" in
success)
# A successful managed workflow says the language analyses
# ran. It does NOT say their security results passed: GitHub
# publishes that verdict as a separate exact-SHA check-run
# owned by github-advanced-security. Requiring only the
# workflow conclusion let a revision with high alerts pass
# this repository's sole required aggregate.
codeql_checks="$(gh_retry gh api --method GET \
"repos/$REPO/commits/$SHA/check-runs?check_name=CodeQL&per_page=100")" || {
rc=$?
if [ "$rc" -eq 22 ]; then exit 1; fi
transport_failed=true
break
}
# Bind by app as well as name and SHA. Another app is free
# to publish a check called CodeQL; it must not answer for
# GitHub's security result. Check-run ids are monotonic, so
# an updated result supersedes an earlier interim row.
codeql_check="$(jq -c '
[.check_runs[]?
| select((.app.slug // "") == "github-advanced-security")]
| max_by(.id) // empty
' <<<"$codeql_checks")"
if [ -z "$codeql_check" ]; then
pending+=("$label (results not reported)")
continue
fi
codeql_status="$(jq -r '.status // ""' <<<"$codeql_check")"
codeql_conclusion="$(jq -r '.conclusion // ""' <<<"$codeql_check")"
if [ "$codeql_status" != "completed" ]; then
pending+=("$label (results $codeql_status)")
else
case "$codeql_conclusion" in
success) passed+=("$label") ;;
# Default setup publishes an interim neutral result while
# one or more configured languages have not reported.
# It is absence of a final result, not a clean verdict.
neutral|"") pending+=("$label (results pending)") ;;
*) failed+=("$label ($codeql_conclusion)") ;;
esac
fi
;;
skipped) passed+=("$label") ;;
*) failed+=("$label ($conclusion)") ;;
esac
elif [ "$label" = "Design Review" ] || [ "$label" = "UX Review" ] || [ "$label" = "First Principles Review" ]; then
# Design, UX and First Principles all BLOCK on a real BLOCK
# verdict. The "gates on BLOCK" step exits 0 for PASS/CONCERNS,
# the Dependabot skip, and any overloaded/throttled/verdict-less
# run, so those never fail the run -- a `failure` conclusion is
# therefore almost always a judged-wrong verdict. One residual:
# a HARD model-step error (OIDC / action crash) has no
# continue-on-error, so it also fails the run and reddens the
# lane. That is rare, deliberate (a review that never ran must
# not read as green), identical to the Design / Opus / GPT
# lanes, and cleared by a one-comment `/ai-review override`
# (current-SHA scoped) -- so treating a `failure` as a blocker
# is safe. (The fork reader above sees one counterpart: the
# fork UX lane completes its check-run as `failure` when an
# attachment download failed for a transport reason; every
# other fork hard error maps to NEUTRAL.)
case "$conclusion" in
# A workflow-level `skipped` is left alone: nothing ran, so no
# verdict is owed for this head and there is no lane to
# re-run. Only a lane that reports having reviewed the head
# is held to having published that review.
skipped) passed+=("$label") ;;
success|neutral)
if [ "$(advisory_slot_unpublished "$label")" = "1" ]; then
pending+=("$label (verdict not published, re-run this lane)")
else
passed+=("$label")
fi ;;
*) failed+=("$label (BLOCK)") ;;
esac
# Security Scope Review is deliberately absent here too, for the
# same reason as in the fork reader above: it is fail-closed, so a
# `failure` conclusion covers both a confirmed regression and a run
# that measured nothing, and only the default branch's plain
# "(failure)" attribution is true of both.
else
case "$conclusion" in
success|neutral) passed+=("$label") ;;
# A workflow-level skip stays NON-BLOCKING -- the paths filters
# skip lanes on purpose, so failing here would redden ordinary
# PRs -- but it is reported as "Not eligible" rather than folded
# into the passed count. Counting it as passed made an absent
# matrix leg indistinguishable from a green one, which is exactly
# how a platform nobody builds for reads as covered. `skipped[]`
# feeds only the summary, never the verdict.
skipped) skipped+=("$label (skipped)") ;;
action_required)
if [ "$FORK" = "true" ]; then
awaiting_approval+=("$label")
else
# On an upstream branch this conclusion is anomalous,
# so retain the failure-class diagnostic.
failed+=("$label ($conclusion)")
fi
;;
*) failed+=("$label ($conclusion)") ;;
esac
fi
fi
done
# A read that still fails after the bounded retries is a TRANSPORT
# failure: this revision's verdict is UNKNOWN, not red. Publish an
# explicit non-terminal "could not be evaluated" verdict (a pending
# status under the existing "readiness: checking" label) instead of
# exiting non-zero, so the publish step still runs and the job never
# leaves a red check for a network blip. This deliberately does NOT
# weaken the gate: pending blocks merge exactly like failure does,
# and only a transport error with NO already-observed blocker takes
# this branch -- a genuine failure or maintainer-approval wait
# recorded by an earlier lane dominates (the array guards below),
# falling through to the normal action-required verdict, so a known
# red is never masked by a later lane's network trouble. Recovery is
# automatic -- the self-heal
# sweep re-fires any pending readiness status older than its
# staleness window, and any later monitored-workflow event recomputes
# sooner.
if [ "$transport_failed" = "true" ] \
&& [ "${#failed[@]}" -eq 0 ] \
&& [ "${#awaiting_approval[@]}" -eq 0 ]; then
# The tiebreaker for every ambiguous case is an asymmetry:
# publishing "pending" can only ever BLOCK a merge, never allow
# one, so pending is the fail-safe write whenever validation
# state is unknown. Concretely: if the SHA already carries a
# BLOCKING verdict (failure/error), this truncated run defers --
# the merge is already held, and overwriting the red with
# pending would only discard its diagnostics and set the sweep
# re-firing. Every other case publishes pending: an existing
# SUCCESS must not be left mergeable when a rerun's validation
# state is unknown (the green may be stale -- that is the unsafe
# direction); an UNREADABLE verdict state gets the same
# treatment, because the worst a pending can do to an unseen
# verdict is block a merge that a re-evaluation will unblock;
# and pending-over-pending / no-status just refresh the
# self-heal sweep's staleness clock.
existing_state="$(bounded gh api "repos/$REPO/commits/$SHA/status" \
--jq '[.statuses[] | select(.context == "PR Readiness")][0].state // empty' \
2>/dev/null)" || existing_state="__unreadable__"
if [ "$existing_state" = "failure" ] || [ "$existing_state" = "error" ]; then
{
echo "state=deferred"
echo "status_state="
} >> "$GITHUB_OUTPUT"
{
echo "### PR Readiness: deferred to an existing blocking verdict"
echo
echo "- PR: #$PR"
echo "- Revision: \`$SHA\`"
echo
echo "This run's evaluation was truncated by a GitHub API"
echo "failure, but the revision already carries a blocking"
echo "readiness verdict (\`$existing_state\`) from another run."
echo "The merge is already held; overwriting that verdict with"
echo "pending would only discard its diagnostics, so nothing"
echo "was published."
} > "$RUNNER_TEMP/pr-readiness-summary.md"
exit 0
fi
# The token below is LOAD-BEARING, not decoration: it is the only
# thing that tells this pending apart from a lane-pending, and
# pr-readiness-sweep.yml matches it to decide that age alone may
# re-fire a recompute here. Keep it FIRST, ahead of the prose,
# because the API caps a description at 140 characters and cuts the
# tail. test_pr_readiness_sweep.py pins the pair.
{
echo "state=could_not_evaluate"
echo "label=readiness: checking"
echo "status_state=pending"
echo "description=$READ_FAILURE_TOKEN Readiness could not be evaluated (transient GitHub API failure); it will be re-evaluated"
} >> "$GITHUB_OUTPUT"
{
echo "### PR Readiness: could not be evaluated"
echo
echo "- PR: #$PR"
echo "- Revision: \`$SHA\`"
echo
echo "A read-only GitHub API call kept failing after bounded"
echo "retries, so this revision's verdict is unknown -- not red."
echo "A non-terminal \`pending\` status was published; the"
echo "self-heal sweep or the next monitored-workflow event will"
echo "re-evaluate it."
} > "$RUNNER_TEMP/pr-readiness-summary.md"
exit 0
fi
# A pull_request_target open/synchronize/reopen/edit run is meant to
# surface a transient "checking" signal while validation spins up.
# But such a run can be runner-queue-delayed by many minutes and
# execute AFTER the workflow_run:completed runs have already published
# the terminal verdict for this same SHA. Seeding this sentinel
# unconditionally would then clobber that decided verdict back to
# "checking" with no further event left to recompute it on the
# unchanged commit -- freezing the status at pending indefinitely.
# Only add it when the live evaluation above still found at least one
# monitored workflow genuinely incomplete; in that case the verdict is
# already pending and the sentinel is merely a human-readable note.
# (A legitimate reviewer re-run arrives via workflow_run and shows up
# as an in_progress workflow above, so it keeps producing "checking"
# through this same real-pending path.)
case "$TRIGGER_EVENT:$TRIGGER_ACTION" in
pull_request_target:opened|pull_request_target:synchronize|pull_request_target:reopened|pull_request_target:edited)
if [ "${#pending[@]}" -gt 0 ]; then
pending+=("validation runs are starting")
fi
;;
esac
# (The long-term "Arbiter — judge from comments" gate was retired: its
# one-way-door / long-term-reversibility lens now lives in Design
# Review. Readiness no longer waits on that retired Arbiter check.)
# Disposition-rule enforcement (issue #6658). A violation is a
# condition waiting cannot fix -- only the author editing or deleting
# the comment can -- so it belongs in `failed`, alongside a red lane.
# `ok=false` means the record set could not be established, which is
# UNKNOWN and therefore `pending`: a transient comments/permission API
# failure must never red the required status (issue #2753's class).
# It is a READ failure, so it is stamped: the sweep may then clear it
# on age, which is the only signal it has, because a recompute whose
# reads succeed resolves it while lane state says nothing about it.
case "${DISPOSITION_OK:-}" in
true)
while IFS= read -r violation; do
[ -n "$violation" ] || continue
failed+=("disposition rule: $violation")
done <<<"${DISPOSITION_VIOLATIONS:-}"
;;
false)
pending+=("disposition records could not be read")
read_failure=true
;;
esac
# Superseded-verdict enforcement (issue #13951). A block that was judged
# at THIS head and then replaced by a clean sample at the same head is a
# condition waiting cannot fix: the replacement already happened, and
# only re-establishing a verdict or a human adjudication resolves it. So
# it belongs in `failed`. Only the dropped-block list gates -- an
# ordinary same-head re-sample is not a defect and must not redden a
# revision. `ok=false` means the question could not be ANSWERED, which
# is UNKNOWN and therefore `pending`, never a red: that covers both an
# unreadable edit history and a head where no lane was examined at all.
# It is a READ failure, so it is stamped like the disposition one.
case "${SUPERSESSION_OK:-}" in
true)
while IFS= read -r lane; do
[ -n "$lane" ] || continue
# Both halves of this text were wrong and each sent the reader
# somewhere the code does not look. `/ai-review override <lane>
# <head>` is an exit for EVERY lane, fork PRs included: the
# lane's re-run replaces the block with the override note, and
# the gate also reads the accepted record itself as adjudicating
# every block the lane raised at this head before it was
# recorded. And what clears the GPT lane by decision
# is the workflow-authored `(all downgraded on adjudication)`
# heading, NOT the `[BLOCK-MERGE-DOWNGRADED]` marker, which sits
# in embedded model output. A new head is the expensive advice: it
# discards every other lane's verdict for this head.
case "$lane" in
GPT) remedy="record a legitimate clear at this head with /ai-review override, or an adjudication pass, whose '(all downgraded on adjudication)' heading is read as cleared; a new head also re-samples it, at the cost of every other lane's verdict for this head" ;;
*) remedy="record a legitimate clear at this head with /ai-review override, whose accepted record adjudicates every block this lane raised at this head before it, on a fork PR too; this lane has no adjudication heading, so the only other exit is a new head, at the cost of every other lane's verdict for this head" ;;
esac
failed+=("superseded verdict: $lane blocked this head in a replaced sample and the body now presented does not - read the comment's stored history; $remedy")
done <<<"${SUPERSESSION_DROPPED:-}"
;;
false)
pending+=("superseded verdicts could not be evaluated")
read_failure=true
;;
esac
blockers=()
waiting=()
if [ "${#failed[@]}" -gt 0 ]; then
blockers=(${failed[@]+"${failed[@]}"})
fi
if [ "${#pending[@]}" -gt 0 ]; then
waiting=(${pending[@]+"${pending[@]}"})
fi
if [ "$DRAFT" = "true" ]; then
waiting+=("pull request is a draft")
fi
if [ "${#blockers[@]}" -gt 0 ] || [ "${#awaiting_approval[@]}" -gt 0 ]; then
state="action_required"
label="readiness: action required"
status_state="failure"
if [ "${#blockers[@]}" -gt 0 ] && [ "${#awaiting_approval[@]}" -gt 0 ]; then
description="${#blockers[@]} blocking readiness item(s); ${#awaiting_approval[@]} awaiting maintainer approval"
elif [ "${#blockers[@]}" -gt 0 ]; then
description="${#blockers[@]} blocking readiness item(s)"
else
description="${#awaiting_approval[@]} workflow(s) awaiting maintainer approval"
if [ "${#passed[@]}" -eq 0 ] && [ "$transport_failed" = "false" ]; then
description="$description; none has run yet"
fi
fi
elif [ "${#waiting[@]}" -gt 0 ]; then
state="checking"
label="readiness: checking"
status_state="pending"
# Name the pending lanes, not just their count. A lane that has
# produced zero runs for an immutable head can never produce one
# without a new commit, so naming it (e.g. "Internal Content Scan
# (not started)") lets a human re-run THAT lane instead of pushing
# an empty commit that forces a full re-review. The commit-status
# description has a hard 140-char API limit (the POST silently
# truncates past it), so the list is capped: names until the budget
# is spent, then "+N more", with the count kept as the prefix. This
# is diagnostic only -- it changes no lane's pass/pending/fail
# classification and no state that resolves or blocks; the same
# `waiting[]` that sets the count sets the names.
count="${#waiting[@]}"
# A read-failure pending reaches the sweep through this generic
# description too, so the token leads this one as well, and the
# names budget below pays for it so the total still fits the cap.
prefix=""
if [ "$read_failure" = "true" ]; then
prefix="$READ_FAILURE_TOKEN "
fi
names_budget=$(( 70 - ${#prefix} ))
names=""
shown=0
for item in ${waiting[@]+"${waiting[@]}"}; do
candidate="$names"
if [ -n "$candidate" ]; then
candidate="$candidate, $item"
else
candidate="$item"
fi
# Reserve room for the count prefix, the "; waiting on " joiner
# and a possible " (+N more)" suffix so the final string cannot
# exceed the 140-char commit-status limit. The count prefix runs
# to ~35 chars and the suffix to ~11, so a 70-char names budget
# leaves margin for both.
if [ "${#candidate}" -gt "$names_budget" ]; then
break
fi
names="$candidate"
shown=$(( shown + 1 ))
done
description="$prefix$count readiness check(s) still pending"
if [ -n "$names" ]; then
description="$description; waiting on $names"
remaining=$(( count - shown ))
if [ "$remaining" -gt 0 ]; then
description="$description (+$remaining more)"
fi
fi
else
# Fork PRs reach this branch too: the AI code reviews run on forks
# via the Stage-2 fork-*-review.yml pipeline and are evaluated above
# from the head SHA's check-runs, so a fully-green fork is fully
# validated (CodeQL is the only ineligible lane, listed as a
# non-blocking "Not eligible" note) and earns "passed" like any
# same-repo PR.
state="passed"
label="readiness: passed"
status_state="success"
description="Eligible automated validation passed for this revision"
fi
# Diagnostic-only: the lane arrays are only ever readable from
# $GITHUB_STEP_SUMMARY, which `gh run view --log` cannot query --
# so a run that publishes a wrong pending/failed verdict (e.g. a
# lane invisible under GITHUB_TOKEN but visible under a user token)
# can only be diagnosed by opening the summary UI by hand (#3550).
# Plain `echo` here lands in the job's own log instead, so
# `gh run view --log | grep 'pr-readiness: lane state'` finds it.
echo "pr-readiness: lane state -- passed=[${passed[*]:-}] pending=[${pending[*]:-}] failed=[${failed[*]:-}] skipped=[${skipped[*]:-}] awaiting_approval=[${awaiting_approval[*]:-}]"
{
echo "state=$state"
echo "label=$label"
echo "status_state=$status_state"
echo "description=$description"
} >> "$GITHUB_OUTPUT"
{
echo "### PR Readiness: ${state//_/ }"
echo
echo "- PR: #$PR"
echo "- Revision: \`$SHA\`"
echo "- Passed components: ${#passed[@]}"
if [ "${#skipped[@]}" -gt 0 ]; then
echo
echo "**Not eligible**"
printf -- "- %s\n" ${skipped[@]+"${skipped[@]}"}
fi
if [ "${#awaiting_approval[@]}" -gt 0 ]; then
echo
echo "**Awaiting maintainer approval**"
printf -- "- %s\n" ${awaiting_approval[@]+"${awaiting_approval[@]}"}
echo
echo "> GitHub created these fork runs but has not executed them."
echo "> A maintainer must approve them; the contributor cannot"
echo "> clear this condition."
fi
if [ "${#blockers[@]}" -gt 0 ]; then
echo
echo "**Blocking**"
printf -- "- %s\n" ${blockers[@]+"${blockers[@]}"}
fi
if [ "${#waiting[@]}" -gt 0 ]; then
echo
echo "**Waiting**"
printf -- "- %s\n" ${waiting[@]+"${waiting[@]}"}
fi
# Reached with transport_failed=true only via the fall-through
# above (an already-observed blocker dominates the pending
# fallback): the arrays stopped filling at the break, so say so --
# otherwise the summary silently reads as a complete evaluation.
if [ "$transport_failed" = "true" ]; then
echo
echo "> Evaluation was truncated by a GitHub API failure before"
echo "> every lane could be read; the verdict above is computed"
echo "> from the lanes read so far and will be recomputed"
echo "> automatically."
fi
} > "$RUNNER_TEMP/pr-readiness-summary.md"
- name: Publish status and label
if: >-
steps.context.outputs.stale != 'true'
&& steps.context.outputs.state == 'OPEN'
&& steps.hold.outputs.held != 'true'
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
REPO: ${{ github.repository }}
PR: ${{ steps.context.outputs.pr }}
SHA: ${{ steps.context.outputs.sha }}
URL: ${{ steps.context.outputs.url }}
TARGET_LABEL: ${{ steps.verdict.outputs.label }}
STATUS_STATE: ${{ steps.verdict.outputs.status_state }}
DESCRIPTION: ${{ steps.verdict.outputs.description }}
FORK: ${{ steps.context.outputs.fork }}
HEAD_REPO: ${{ steps.context.outputs.head_repo }}
HEAD_REF: ${{ steps.context.outputs.head_ref }}
run: |
set -euo pipefail
source "$RUNNER_TEMP/gh-retry.sh"
# A truncated evaluation that deferred to an existing terminal
# verdict emits no status_state: there is nothing safe to publish.
if [ -z "$STATUS_STATE" ]; then
echo "Evaluation deferred to an existing verdict on $SHA; publishing nothing."
exit 0
fi
current_sha="$(gh_retry gh pr view "$PR" --repo "$REPO" --json headRefOid --jq '.headRefOid')"
if [ "$current_sha" != "$SHA" ]; then
echo "Ignoring publish for stale revision $SHA; PR #$PR is now at $current_sha."
exit 0
fi
# Re-check before publishing a SUCCESS. The verdict was scored from a
# runs page read before the checkout and the disposition gate, and a
# run in an isolated group (pull_request_target, workflow_dispatch)
# is never cancelled by a lane re-run's own readiness run. A lane
# that started since would otherwise be published over as success --
# the one write that lets armed auto-merge land while a lane runs,
# and the write a re-run's hold cannot defend against: the hold lands
# in seconds and this publish lands after it. One read, on every
# success; any open monitored lane downgrades this publish to
# pending, and a failed read does too (stamped, so the self-heal
# sweep re-fires it).
#
# An `in_progress`-triggered run never reaches this step: it publishes
# its hold and stops (see "Hold the verdict while a monitored lane
# runs"), so no event-shape downgrade is needed here.
#
# The runs-page filter below matches on this PR's own head repository
# and branch, which a fork head's runs DO carry (a fork PR's CI run
# lists the fork as head_repository and its branch as head_branch),
# so this read answers for a fork's CI, Fast Gate, Build, Code Review
# and content-scan re-runs exactly as it does for a same-repo PR. What
# it cannot see is the seven Stage-2 fork review lanes, which run from
# the default branch and report as check-runs on the head SHA; those
# get their own read further down.
if [ "$STATUS_STATE" = "success" ]; then
# The page is kept: the fork arm below reads Fast Gate's newest run
# from it rather than spending a second request.
runs_page=""
if runs_page="$(gh_retry gh api --method GET --paginate \
"repos/$REPO/actions/runs?event=pull_request&head_sha=$SHA&per_page=100")" \
&& recheck="$(jq -rs --arg repo "$HEAD_REPO" --arg ref "$HEAD_REF" --arg lanes "$MONITORED_LANES" '
($lanes | split(" ") | map(select(. != "")) | map(".github/workflows/" + .)) as $l
| [.[].workflow_runs[]
| select((.path | IN($l[]))
and (.head_repository.full_name // "") == $repo
and (.head_branch // "") == $ref)]
# Newest run per lane, collapsed exactly as the verdict step
# collapses it: max id, and a cancelled max-id run yields to a
# LATER-STARTED sibling (fork approval can start runs out of id
# order, so the max-id twin is the one that got cancelled).
# Reading a lane differently here than the evaluation did would
# red a verdict it scored green, deterministically, on every
# recompute.
| group_by(.path)
| map(
(max_by(.id)) as $n
| if ($n.conclusion // "") == "cancelled" then
([.[] | select(.id != $n.id and .run_started_at != null
and .run_started_at >= $n.run_started_at)]
| max_by([.run_started_at, .conclusion != "cancelled"])) // $n
else $n end)
| (map(select(.status == "completed"
and ((.conclusion // "") | IN("success","neutral","skipped") | not))
| (.name // .path)) | join(", ")) as $red
| (map(select(.status != "completed") | (.name // .path)) | join(", ")) as $open
| "\($red)|\($open)"
' <<<"$runs_page")"; then
IFS='|' read -r red_lanes open_lanes <<<"$recheck"
# A red lane publishes red: a pending here would overwrite the
# failure that lane's own evaluation already published.
if [ -n "$red_lanes" ]; then
echo "pr-readiness: publish re-check -- lane(s) red since the evaluation: $red_lanes"
STATUS_STATE=failure
TARGET_LABEL="readiness: action required"
DESCRIPTION="blocking readiness item(s) since evaluation: $red_lanes"
elif [ -n "$open_lanes" ]; then
echo "pr-readiness: publish re-check -- lane(s) started since the evaluation: $open_lanes"
STATUS_STATE=pending
TARGET_LABEL="readiness: checking"
DESCRIPTION="readiness check(s) changed since evaluation; re-checking $open_lanes"
fi
else
STATUS_STATE=pending
TARGET_LABEL="readiness: checking"
DESCRIPTION="[read-failed] Readiness could not be re-checked before publishing success; it will be re-evaluated"
fi
# CodeQL is a `dynamic` run, so the re-check above -- filtered to
# `event=pull_request` and MONITORED_LANES -- is structurally blind
# to it. Without this, a CodeQL re-run had no
# way to hold the publish: the delayed publisher saw only green
# pull-request lanes and wrote success over the pending that re-run's
# own readiness run had published, opening a merge window while the
# code-scanning verdict for the revision was in flight. Read on the
# same two reads the verdict step scores CodeQL from, and only while
# this publish is still a success. An empty dynamic page holds
# nothing: CodeQL either does not apply to this base or has not
# started, and in the latter case the verdict step said pending.
# Each conclusion is mapped to the verdict step's own reading of it,
# arm for arm: `success` goes on to the result check-run, `skipped`
# is that step's `passed` (default setup did not scan this base),
# and every other completed conclusion is its `failed`. Reading one
# of them differently here makes the re-check disagree with the
# evaluation it is re-checking -- either restoring a success that
# evaluation rejected, or blocking a verdict it passed. A fork head
# cannot run the managed default-setup CodeQL (the verdict step
# lists it skipped), so this read is same-repo only.
if [ "$STATUS_STATE" = "success" ] && [ "${FORK:-}" = "false" ]; then
if codeql_state="$(gh_retry gh api --method GET \
"repos/$REPO/actions/runs?event=dynamic&head_sha=$SHA&per_page=100" \
| jq -r '
[.workflow_runs[]
| select(.path == "dynamic/github-code-scanning/codeql")]
| max_by(.id) as $r
| if $r == null then "none|"
elif ($r.status // "") != "completed" then "open|CodeQL (\($r.status // "queued"))"
elif ($r.conclusion // "") == "success" then "results|"
elif ($r.conclusion // "") == "skipped" then "none|"
else "red|CodeQL (\($r.conclusion))" end
')"; then
IFS='|' read -r codeql_verdict codeql_lane <<<"$codeql_state"
# A successful managed workflow says the analyses ran, not that
# their results passed: that verdict is a separate exact-SHA
# check-run owned by github-advanced-security, and it is the one
# a re-scan re-opens. Bound by app as well as name, as the
# verdict step does, so another app's `CodeQL` cannot answer.
if [ "$codeql_verdict" = "results" ]; then
if codeql_state="$(gh_retry gh api --method GET \
"repos/$REPO/commits/$SHA/check-runs?check_name=CodeQL&per_page=100" \
| jq -r '
[.check_runs[]?
| select((.app.slug // "") == "github-advanced-security")]
| max_by(.id) as $c
| if $c == null then "open|CodeQL (results not reported)"
elif ($c.status // "") != "completed" then "open|CodeQL (results \($c.status))"
elif (($c.conclusion // "") | IN("success")) then "none|"
elif (($c.conclusion // "") | IN("neutral","")) then "open|CodeQL (results pending)"
else "red|CodeQL (\($c.conclusion))" end
')"; then
IFS='|' read -r codeql_verdict codeql_lane <<<"$codeql_state"
else
codeql_verdict=unreadable
fi
fi
case "$codeql_verdict" in
red)
echo "pr-readiness: publish re-check -- $codeql_lane since the evaluation"
STATUS_STATE=failure
TARGET_LABEL="readiness: action required"
DESCRIPTION="blocking readiness item(s) since evaluation: $codeql_lane"
;;
open)
echo "pr-readiness: publish re-check -- $codeql_lane since the evaluation"
STATUS_STATE=pending
TARGET_LABEL="readiness: checking"
DESCRIPTION="readiness check(s) changed since evaluation; re-checking $codeql_lane"
;;
unreadable)
STATUS_STATE=pending
TARGET_LABEL="readiness: checking"
DESCRIPTION="[read-failed] Readiness could not be re-checked before publishing success; it will be re-evaluated"
;;
esac
else
STATUS_STATE=pending
TARGET_LABEL="readiness: checking"
DESCRIPTION="[read-failed] Readiness could not be re-checked before publishing success; it will be re-evaluated"
fi
fi
# The seven Stage-2 fork review lanes run from the default branch,
# so the runs page above never shows them on a fork head; the
# verdict step reads them from the head SHA's check-runs, and so
# does this re-check, bound the same way: each lane stamps its row
# with `<prefix><PR>-<Fast Gate run id>-<attempt>`, and only a row
# naming THIS pull request AND the newest Fast Gate run + attempt
# for this head is current. The attempt is what makes a Fast Gate
# re-run visible: the previous attempt's seven rows still sit on
# the SHA, completed and green, and a match on PR alone would read
# them as the verdict while the re-run's Stage-2 lanes have not yet
# posted theirs. A lane with no current-attempt row has not
# reported and holds the publish, as the verdict step's own
# "(not started)" does; a current row still running holds it; a
# current row completed red is a judged-wrong verdict -- each fork
# lane fails its check ONLY on a real BLOCK, an errored or
# throttled run resolves neutral -- and publishes red. Fast Gate's
# newest run is read from the page already fetched above, with the
# verdict step's own max-id collapse. One request, on a fork
# success only.
if [ "$STATUS_STATE" = "success" ] && [ "${FORK:-}" = "true" ]; then
fast_gate="$(jq -rs --arg repo "$HEAD_REPO" --arg ref "$HEAD_REF" '
[.[].workflow_runs[]
| select(.path == ".github/workflows/fast-gate.yml"
and (.head_repository.full_name // "") == $repo
and (.head_branch // "") == $ref)]
| max_by(.id) | if . == null then "" else "\(.id)-\(.run_attempt // 1)" end
' <<<"$runs_page")"
if [ -z "$fast_gate" ]; then
# No Fast Gate run on this head means no Stage-2 lane can have
# reported; the verdict step would have said pending.
echo "pr-readiness: publish re-check -- Fast Gate has no run on this fork head; Stage-2 lanes cannot have reported"
STATUS_STATE=pending
TARGET_LABEL="readiness: checking"
DESCRIPTION="readiness check(s) changed since evaluation; re-checking Fast Gate"
elif fork_state="$(gh_retry gh api --method GET --paginate \
"repos/$REPO/commits/$SHA/check-runs?per_page=100" \
| jq -rs --arg pr "$PR" --arg fg "$fast_gate" \
--argjson legacy "$LEGACY_LANE_NAMES" '
($legacy | with_entries({key: .value, value: .key})) as $old_name
| {"Internal Content Scan": "internal-content-scan-pr-",
"Opus 5.5 Review": "opus-pr-",
"GPT 6.1 Review": "gpt-pr-",
"Design Review": "design-pr-",
"UX Review": "ux-pr-",
"First Principles Review": "first-principles-pr-",
"Security Scope Review": "scope-pr-"} as $lanes
| [.[].check_runs[]] as $rows
| [$lanes | to_entries[]
| .key as $name | (.value + $pr + "-" + $fg) as $want
| ($rows | map(select((.external_id // "") == $want))) as $bound
# A pre-rename row answers only when no current-name row
# is bound; see LEGACY_LANE_NAMES.
| (($bound | map(select(.name == $name)) | max_by(.id))
// ($bound | map(select(.name == $old_name[$name])) | max_by(.id))) as $row
| {name: $name, row: $row}]
| (map(select(.row != null and .row.status == "completed"
and ((.row.conclusion // "")
| IN("failure","timed_out","cancelled",
"action_required","stale","startup_failure"))))
| map(.name)) as $red
| (map(select(.row == null) | .name + " (not started)")
+ map(select(.row != null and .row.status != "completed") | .name)) as $open
| "\($red | join(", "))|\($open | join(", "))"
')"; then
IFS='|' read -r fork_red fork_open <<<"$fork_state"
if [ -n "$fork_red" ]; then
echo "pr-readiness: publish re-check -- fork lane(s) red since the evaluation: $fork_red"
STATUS_STATE=failure
TARGET_LABEL="readiness: action required"
DESCRIPTION="blocking readiness item(s) since evaluation: $fork_red"
elif [ -n "$fork_open" ]; then
echo "pr-readiness: publish re-check -- fork lane(s) not current since the evaluation: $fork_open"
STATUS_STATE=pending
TARGET_LABEL="readiness: checking"
DESCRIPTION="readiness check(s) changed since evaluation; re-checking $fork_open"
fi
else
STATUS_STATE=pending
TARGET_LABEL="readiness: checking"
DESCRIPTION="[read-failed] Readiness could not be re-checked before publishing success; it will be re-evaluated"
fi
fi
DESCRIPTION="${DESCRIPTION:0:140}"
fi
# The commit status IS the verdict, and `PR Readiness` is the only
# required status check on protected branches, so it is published
# FIRST and on its own. The labels below are advisory decoration; a
# cosmetic label call must never be in a position to decide whether
# the required check reports anything at all.
#
# This POST is deliberately NOT retried. A commit status is
# last-write-wins per SHA+context, GitHub offers no conditional
# write, and two runs can evaluate the SAME revision concurrently
# (see the concurrency comment above) -- so ANY retry races a
# concurrent run's newer verdict: between whatever guard read a
# retry could take and its re-POST, another run can publish, and
# the re-POST would republish this run's verdict over it (a stale
# green over a fresh red, or a pending over a terminal). Three
# successive guard designs ("any status exists", compare-and-swap
# on the status id, fail-safe reads) each narrowed but could not
# close that window -- the API gives no atomic primitive to close
# it with. A failed POST therefore fails the step loud instead:
# the run shows red with an actionable error, a re-run (or the
# next monitored-workflow event) re-evaluates and republishes,
# and no verdict is ever silently overwritten. The POST is
# bounded (gh applies no HTTP timeout of its own).
jq -n \
--arg state "$STATUS_STATE" \
--arg target_url "$URL" \
--arg description "$DESCRIPTION" \
'{
state: $state,
target_url: $target_url,
description: $description,
context: "PR Readiness"
}' > "$RUNNER_TEMP/readiness-status.json"
if ! bounded gh api --method POST "repos/$REPO/statuses/$SHA" \
--input "$RUNNER_TEMP/readiness-status.json" >/dev/null; then
echo "::error::Failed to publish the readiness status for $SHA." \
"The POST is not retried (a retry can overwrite a concurrent" \
"run's newer verdict); re-run this workflow to re-evaluate" \
"and republish."
exit 1
fi
# Two runs can evaluate the SAME revision concurrently: each
# pull_request_target run gets its own concurrency group keyed on
# run_id, which keeps a superseded run from showing as cancelled and
# leaves two runs on one revision racing each other's labels. Both
# races below end in the state this run wanted, so neither is an
# error -- but nothing wider than those two is tolerated.
drop_label() {
local label="$1" encoded out
encoded="$(jq -rn --arg value "$label" '$value | @uri')"
if out="$(gh api --method DELETE \
"repos/$REPO/issues/$PR/labels/$encoded" 2>&1)"; then
return 0
fi
if grep -q 'HTTP 404' <<<"$out"; then
echo "Label '$label' was already removed by a concurrent run."
return 0
fi
printf '%s\n' "$out" >&2
return 1
}
ensure_label() {
local name="$1" color="$2" description="$3" out
if out="$(gh label create "$name" --repo "$REPO" \
--color "$color" --description "$description" 2>&1)"; then
return 0
fi
if grep -qi 'already exists' <<<"$out"; then
echo "Label '$name' was already created by a concurrent run."
return 0
fi
printf '%s\n' "$out" >&2
return 1
}
label_specs=(
"readiness: checking|BFDADC|Automated validation is still running"
"readiness: action required|D73A4A|A blocking check or review needs attention"
"readiness: passed|0E8A16|Eligible automated validation passed for the current revision"
)
existing_labels="$(gh_retry gh label list --repo "$REPO" --limit 200 --json name --jq '.[].name')"
for spec in ${label_specs[@]+"${label_specs[@]}"}; do
name="${spec%%|*}"
rest="${spec#*|}"
color="${rest%%|*}"
description="${rest#*|}"
if ! grep -Fxq "$name" <<<"$existing_labels"; then
ensure_label "$name" "$color" "$description"
fi
done
current="$(gh_retry gh api "repos/$REPO/issues/$PR/labels" --paginate --jq '.[].name')"
for label in \
"readiness: checking" \
"readiness: action required" \
"readiness: maintainer review" \
"readiness: passed"; do
if [ "$label" != "$TARGET_LABEL" ] && grep -Fxq "$label" <<<"$current"; then
drop_label "$label"
fi
done
if ! grep -Fxq "$TARGET_LABEL" <<<"$current"; then
# Not `--arg label`: `label` is a jq keyword that jq 1.6 rejects.
jq -n --arg name "$TARGET_LABEL" '{labels: [$name]}' \
| gh api --method POST "repos/$REPO/issues/$PR/labels" \
--input - >/dev/null
fi
cat "$RUNNER_TEMP/pr-readiness-summary.md" >> "$GITHUB_STEP_SUMMARY"
# One line for measuring this workflow's draw on the shared pool.
# `GET /rate_limit` is not counted against the limit it reports.
- name: Log the remaining REST budget
if: always()
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: |
timeout 60 gh api rate_limit --jq '.resources.core
| "pr-readiness: core rate limit -- \(.remaining)/\(.limit) remaining, resets \(.reset | todate)"' \
|| echo "pr-readiness: core rate limit could not be read"