Skip to content

fix(release): reclaim docker disk and gate the data root before the sandbox image build (#13479) - #13481

Open
qwen-code-dev-bot wants to merge 11 commits into
mainfrom
autofix/issue-13479
Open

qwen-code-dev-bot wants to merge 11 commits into
mainfrom
autofix/issue-13479

Conversation

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator

What this PR does

Hardens the nightly release's Docker integration lane against runner disk exhaustion in two ways. Before the lane builds the sandbox image, it now also prunes BuildKit build cache (the previous cleanup only pruned labelled images, which cannot reclaim cache), using the same 24-hour age filter as the daily host sweep. Right after that reclaim, it re-checks free space — this time on the Docker data root filesystem, at a build-sized floor of 8 GiB, reusing the existing disk-floor script — and fails the job fast with a legible annotation when a host is still saturated.

Both behaviors are scoped to the branch that actually builds the image; when a concurrent run already built the image for the same revision, nothing changes. The change lives in the extracted release script that rides github.workflow_sha, so it takes effect on the next nightly without touching any release in flight.

Why it's needed

The 2026-10-05 nightly release (#13479) died 24 minutes into the Docker test step: the self-hosted runner hit ENOSPC and the runner worker process itself crashed writing its own diagnostic log. The job-start disk floor gate had passed — it runs before the build by up to an hour, and its 2 GiB floor is far below what the sandbox image build (a full monorepo install and bundle inside Docker) needs. Because the worker crashed, the step ended with no conclusion, even the always() container-cleanup step never ran, and the auto-filed issue carried no actionable signal. This is the same ENOSPC class as #10035; the job-start gates added in #11972/#12769 simply do not cover the Docker lane's heaviest step. With this PR, a host full of reclaimable Docker garbage gets reclaimed and the build proceeds; a genuinely saturated host fails in seconds with a clear ::error:: so a re-run lands on an instance with headroom, instead of crashing the runner mid-build.

Reviewer Test Plan

How to verify

Read the diff of the extracted release Docker script: inside the image-missing branch, a docker builder prune --all --force --filter 'until=24h' now follows the existing labelled-image prune, and a guarded call to check-disk-floor.sh against the Docker data root (8 GiB floor, overridable via DISK_FLOOR_MIN_FREE_KB) stands between the prunes and the image build. Then confirm the new script tests: npx vitest run --config ./scripts/tests/vitest.config.ts release-workflow — the suite drives the real script against stub docker/npm/npx/git/node binaries and asserts the build happens only after prune → cache-prune → floor gate in that order, that a tripped gate stops the build, and that an already-present image skips both. Each guard was mutation-probed: deleting the gate block, degrading it to warn-only, or deleting the cache-prune line turns named tests red.

Evidence (Before & After)

N/A (CI shell script; no user-visible surface). Before: run 37374675168's docker lane shows a runner-level System.IO.IOException: No space left on device annotation with no failed step and no cleanup. After: the same saturation surfaces as a fast, annotated gate failure before the build starts, or is reclaimed away entirely.

Tested on

OS Status
🍏 macOS ⚠️ not tested
🪟 Windows ⚠️ not tested (new execution tests skip on win32, matching the file's existing bash-driven tests)
🐧 Linux ✅ tested

Environment (optional)

Unit/script-level only: ShellCheck 0.11.0 (the version pinned by scripts/lint.js) clean on the changed script; full scripts vitest suite green apart from two pre-existing environment-dependent failures (verify-capture, web-shell-publish-artifacts) that fail identically on the base tree; npm run build, npm run typecheck, npm run lint all pass. The Docker lane itself was not re-run — no Docker daemon in the authoring environment.

Risk & Scope

Linked Issues

Fixes #13479

中文说明

本 PR 做了什么

从两个方面加固 nightly 发布的 Docker 集成通道,抵御 runner 磁盘耗尽。其一,在该通道构建沙箱镜像之前,现在会额外裁剪 BuildKit 构建缓存(此前的清理只裁剪带标签镜像,无法回收缓存),时长过滤器与每日主机清理同为 24 小时。其二,回收之后立即重新检查剩余空间——这次检查的是 Docker 数据根目录所在的文件系统,下限按构建体量设为 8 GiB,复用现有的磁盘下限脚本——若主机仍处于饱和状态,则快速失败并输出清晰的注解。

这两个行为都限定在真正执行镜像构建的分支内;当并发运行已为同一修订版本构建好镜像时,行为完全不变。改动位于随 github.workflow_sha 检出的发布脚本中,因此对下一次 nightly 自动生效,不影响任何进行中的发布。

为什么需要

2026-10-05 的 nightly 发布(#13479)在 Docker 测试步骤进行到约 24 分钟时失败:自托管 runner 磁盘写满(ENOSPC),runner worker 进程在写自身诊断日志时崩溃。作业开始时的磁盘下限检查是通过了的——它比构建早最多一小时,且 2 GiB 的下限远低于沙箱镜像构建(在 Docker 内做完整 monorepo 安装与打包)的实际需求。由于 worker 崩溃,该步骤没有任何结论,连 always() 条件的容器清理步骤都未能运行,自动提交的 issue 也没有任何可操作的信号。这与 #10035 是同一类 ENOSPC 问题;#11972/#12769 增加的作业级下限门恰恰没有覆盖 Docker 通道最重的步骤。有了本 PR,充满可回收 Docker 垃圾的主机会被回收后继续构建;真正饱和的主机则会在几秒内带着清晰的 ::error:: 失败,使重跑能落到有余量的实例上,而不是让 runner 在构建中途崩溃。

评审者测试计划

如何验证

审阅提取出的发布 Docker 脚本的 diff:在镜像不存在的分支内,docker builder prune --all --force --filter 'until=24h' 现在紧跟现有的带标签镜像裁剪,而对 Docker 数据根目录调用 check-disk-floor.sh 的受保护检查(8 GiB 下限,可通过 DISK_FLOOR_MIN_FREE_KB 覆盖)位于裁剪与镜像构建之间。然后确认新增的脚本测试:npx vitest run --config ./scripts/tests/vitest.config.ts release-workflow——该套件用 stub 的 docker/npm/npx/git/node 二进制驱动真实脚本,断言构建只发生在镜像裁剪 → 缓存裁剪 → 下限门之后且顺序固定,断言下限门触发时构建不会执行,断言镜像已存在时两者都跳过。每个守卫都经过变异探测:删除门检查代码块、将其降级为仅告警、或删除缓存裁剪行,都会使指定测试变红。

证据(前后对比)

N/A(CI shell 脚本,无用户可见界面)。修复前:运行 37374675168 的 Docker 通道只有 runner 级的 System.IO.IOException: No space left on device 注解,没有失败步骤,也没有清理。修复后:同样的饱和状态要么被回收消解,要么表现为构建开始前的快速、带注解的门失败。

测试平台

系统 状态
🍏 macOS ⚠️ 未测试
🪟 Windows ⚠️ 未测试(新增执行测试在 win32 上跳过,与该文件既有 bash 驱动测试一致)
🐧 Linux ✅ 已测试

环境(可选)

仅单元/脚本级验证:scripts/lint.js 钉住的 ShellCheck 0.11.0 对被改脚本无告警;scripts vitest 全套件除两个既有环境相关失败(verify-capture、web-shell-publish-artifacts,在基线树上同样失败)外全绿;npm run build、npm run typecheck、npm run lint 全部通过。Docker 通道本身未重跑——编写环境没有 Docker daemon。

风险与范围

关联 Issue

Fixes #13479

…andbox image build (#13479)

The nightly release's docker lane died 24 minutes into its test step when
the self-hosted runner hit ENOSPC and the runner worker crashed, after
passing the job-start disk floor gate (run 37374675168). The gate predates
the sandbox image build by up to an hour, and the lane's pre-build prune
only removes labelled images — BuildKit cache and dangling layers on the
shared pool's persistent daemon store are out of its reach.

Prune BuildKit cache alongside the labelled-image prune, then re-gate the
docker data root filesystem at a build-sized floor (8 GiB) immediately
before the build. A host with reclaimable docker garbage now gets reclaimed
and builds; a genuinely saturated host fails fast with a legible ::error::
so a re-run lands on an instance with headroom, instead of the runner
worker crashing mid-build and killing even the cleanup steps.

Co-authored-by: Qwen-Coder <[email protected]>
@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

Autofix report for #13479 — Release Failed for v0.25.0-nightly.20261005.69d5db2ff2

Diagnosis

The failing job was Integration Tests (Docker) on self-hosted runner ecs-qwen-hk4-26. The job's GitHub annotation shows the runner worker itself crashed:

System.IO.IOException: No space left on device : '/home/github-runner/actions-runner-26/_diag/Worker_20261005-211802-utc.log'

Timeline reconstructed from the public API: the job-start Disk floor gate (self-hosted) passed at 21:18:52Z (≥ 2 GiB free), the Run Docker Integration Tests step started at 21:22:37Z, and the runner died ~24 minutes into it (job ended 21:46:51Z) — mid-step, with no step conclusion and without even the always() container-cleanup step running. Every other release lane, including Integration Tests (No Sandbox) running the same tests on the same pool, passed.

Root cause: the docker lane's heavy step — the sandbox image build, a full monorepo install + bundle inside a two-stage Dockerfile — runs long after the job-start disk floor gate, on a shared pool whose docker daemon store persists across jobs. The lane's pre-build cleanup only pruned labelled images older than 24h; BuildKit's intermediate install/build layers (the build's largest disk consumer) and dangling images were outside its reach. The daily host sweep (qwen-docker-cleanup.timer, 02:30 UTC) bounds the store, but the release builds at ~21:20 UTC, at the peak of the day's accumulation. This is the same ENOSPC class as #10035 (which added the job-start gates via #10394/#11972/#12769) — tonight's run shows the gate's coverage ends before the docker lane's heaviest step begins.

No product code defect is involved; the previous four nightly releases passed this lane, and the two September release failures were ordinary test failures with completed steps, a different signature.

Fix

One focused change in .github/scripts/run-release-docker-integration.sh, inside the existing image-missing (i.e. about-to-build) branch, after the lock acquisition and the existing labelled-image prune:

  1. docker builder prune --all --force --filter 'until=24h' — reclaim BuildKit cache, which image pruning cannot touch, using the same age filter as the daily host sweep. Warn-only on failure, like the existing prune.
  2. A pre-build disk floor gate on the Docker data root (docker info --format '{{.DockerRootDir}}'), reusing the existing check-disk-floor.sh with an 8 GiB floor (build-sized: cold builder stage + final image, with margin; overridable via DISK_FLOOR_MIN_FREE_KB). A breached floor now fails the job fast with a legible ::error:: — the established "a re-run lands on an instance with headroom" pattern from fix(release): gate pool-routed validation jobs on a disk floor #11972 — instead of the runner worker crashing mid-build and killing the cleanup steps with it.

Both pieces are scoped to the build branch: when the image already exists (built by a concurrent run at the same revision) nothing changes. Because release workflow scripts are checked out from github.workflow_sha, the fix rides main and applies to the next nightly automatically.

Coverage: four new tests in scripts/tests/release-workflow.test.js — a text/ordering pin, plus three execution tests that drive the real script against stub docker/npm/npx/git/node binaries: (a) build happens only after image prune → cache prune → floor gate, in that order; (b) a tripped gate stops the build and the tests; (c) an already-present image skips reclaim and gate entirely.

Environment note

This runner has no Docker daemon, so the release docker lane itself could not be re-run locally; verification is by the stub-harness execution tests above, shellcheck, and the full scripts suite. The mutation probes below confirm each new guard has a failing test without it.

Verification

  • bash -n .github/scripts/run-release-docker-integration.sh — passed.
  • node scripts/lint.js --shellcheck — could not run in this sandbox: the lane stages its file list with file --mime-type, and the file binary is not installed here. Ran the lane's pinned ShellCheck 0.11.0 (downloaded from the URL in scripts/lint.js) directly instead: clean on the changed script and on check-disk-floor.sh.
  • npx vitest run --config ./scripts/tests/vitest.config.ts release-workflow — 71 passed (67 pre-existing + 4 new).
  • npx vitest run --config ./scripts/tests/vitest.config.ts (full scripts suite) — 2800 passed, 12 failed in verify-capture.test.js (terminal rendering) and web-shell-publish-artifacts.test.js (npm pack output format); stashing this change and re-running the base tree fails the identical 12, so they are pre-existing environment issues unrelated to this change.
  • Mutation probes (remove guard → expect red → restore → green): (1) gate block deleted → 3 tests failed; (2) gate degraded to warn-only || true → gate-failure test failed; (3) docker builder prune line deleted → 2 tests failed. All restored and re-verified green.
  • node --test .github/scripts/check-disk-floor.test.mjs .github/scripts/ecs-runner/qwen-docker-cleanup.test.mjs — 10/10 passed.
  • npm run build — passed.
  • npm run typecheck — passed.
  • npm run lint — passed.
  • Not run: the docker integration lane itself (no Docker daemon available in this environment).
中文说明

#13479 自动修复报告 — v0.25.0-nightly.20261005.69d5db2ff2 发布失败

诊断

失败的作业是运行在自托管 runner ecs-qwen-hk4-26 上的 Integration Tests (Docker)。该作业在 GitHub 上的注解显示 runner worker 进程自身崩溃了:

System.IO.IOException: No space left on device : '/home/github-runner/actions-runner-26/_diag/Worker_20261005-211802-utc.log'

根据公开 API 重建的时间线:作业开始时的 Disk floor gate (self-hosted) 在 21:18:52Z 通过(剩余空间 ≥ 2 GiB),Run Docker Integration Tests 步骤在 21:22:37Z 开始,约 24 分钟后 runner 死亡(作业于 21:46:51Z 结束)——步骤中途崩溃,没有任何步骤结论,连带 always() 条件的容器清理步骤也没有运行。其他所有发布通道全部通过,包括在同一主机池上运行相同测试的 Integration Tests (No Sandbox)。

根本原因:Docker 通道最重的步骤——沙箱镜像构建(在两阶段 Dockerfile 内做完整 monorepo 安装 + 打包)——发生在作业级磁盘下限检查之后很久,且运行在 daemon 存储跨作业持久化的共享主机池上。该通道构建前的清理只裁剪 24 小时前的带标签镜像;BuildKit 的中间安装/构建层(构建最大的磁盘消耗者)和悬空镜像都不在其覆盖范围内。每日主机清理(qwen-docker-cleanup.timer,UTC 02:30)虽然对存储设有上限,但发布在约 UTC 21:20 构建,正值全天积累的峰值。这与 #10035 是同一类 ENOSPC 问题(该问题通过 #10394/#11972/#12769 引入了作业级下限门)——今晚的运行表明该门的覆盖范围在 Docker 通道最重的步骤开始之前就结束了。

不涉及任何产品代码缺陷:前四次 nightly 发布都通过了该通道;九月的两次发布失败是普通测试失败(步骤正常完成),签名不同。

修复

仅修改 .github/scripts/run-release-docker-integration.sh 一处:在现有的"镜像不存在(即即将构建)"分支内、获取锁与现有带标签镜像裁剪之后:

  1. docker builder prune --all --force --filter 'until=24h'——回收镜像裁剪无法触及的 BuildKit 缓存,使用与每日主机清理相同的时长过滤器。失败时仅告警,与现有裁剪一致。
  2. 在构建前对 Docker 数据根目录(docker info --format '{{.DockerRootDir}}')做磁盘下限检查,复用现有的 check-disk-floor.sh,下限为 8 GiB(按构建体量设定:冷构建阶段 + 最终镜像,留有余量;可通过 DISK_FLOOR_MIN_FREE_KB 覆盖)。下限被突破时,作业会快速失败并输出清晰的 ::error::——沿用 fix(release): gate pool-routed validation jobs on a disk floor #11972 确立的"重跑会落到有余量的实例上"模式——而不是让 runner worker 在构建中途崩溃并连带杀掉清理步骤。

两处改动都限定在构建分支内:当镜像已存在(由同一修订版本的并发运行构建)时行为完全不变。由于发布工作流脚本是从 github.workflow_sha 检出的,该修复随 main 走,会自动生效于下一次 nightly。

测试覆盖:在 scripts/tests/release-workflow.test.js 中新增四个测试——一个文本/顺序钉住测试,外加三个用 stub 的 docker/npm/npx/git/node 二进制驱动真实脚本的执行测试:(a) 构建只发生在镜像裁剪 → 缓存裁剪 → 磁盘下限门之后,且顺序固定;(b) 下限门触发时构建与测试都不会执行;(c) 镜像已存在时完全跳过回收与门检查。

环境说明

本 runner 没有 Docker daemon,无法在本地重跑发布 Docker 通道;验证方式为上述 stub 执行测试、shellcheck 以及完整 scripts 测试套件。下面的变异探测确认每个新增守卫在没有它时都有测试变红。

验证

  • bash -n .github/scripts/run-release-docker-integration.sh——通过。
  • node scripts/lint.js --shellcheck——无法在本沙箱中运行:该通道用 file --mime-type 生成文件列表,而此环境未安装 file 二进制。改为直接运行该通道钉住的 ShellCheck 0.11.0(从 scripts/lint.js 中的 URL 下载):被改脚本与 check-disk-floor.sh 均无告警。
  • npx vitest run --config ./scripts/tests/vitest.config.ts release-workflow——71 个通过(67 个既有 + 4 个新增)。
  • npx vitest run --config ./scripts/tests/vitest.config.ts(完整 scripts 套件)——2800 个通过,12 个失败,位于 verify-capture.test.js(终端渲染)与 web-shell-publish-artifacts.test.js(npm pack 输出格式);将本次改动 stash 后在基线树上重跑,失败的恰是同样 12 个,属于与本改动无关的既有环境问题。
  • 变异探测(移除守卫 → 预期变红 → 恢复 → 复绿):(1) 删除门检查代码块 → 3 个测试失败;(2) 将门降级为仅告警(|| true)→ 门失败测试变红;(3) 删除 docker builder prune 行 → 2 个测试失败。全部恢复后重新验证为绿。
  • node --test .github/scripts/check-disk-floor.test.mjs .github/scripts/ecs-runner/qwen-docker-cleanup.test.mjs——10/10 通过。
  • npm run build——通过。
  • npm run typecheck——通过。
  • npm run lint——通过。
  • 未运行:Docker 集成通道本身(本环境无 Docker daemon)。

🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.25.0

@qwen-code-dev-bot

qwen-code-dev-bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator Author

✅ AutoFix round 6 finished — view run. See this round's report below.

中文说明

✅ AutoFix 第 6 轮已完成 —— 查看运行。本轮报告见下方。

…#13479)

Address the round-1 review of #13481: add the missing dangling-image
prune, bound the BuildKit cache prune with timeout 20m inside the host
build mutex, scope the floor gate to self-hosted runners like every
other check-disk-floor.sh call site, warn on every gate-skip path,
re-gate after the build (the #13479 death landed in the vitest phase),
and extend the harness with a self-hosted lock-protocol case, a strict
docker info stub, and prune/gate failure cases.

Co-authored-by: Qwen-Coder <[email protected]>
@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🤖 Addressed the latest review feedback (round 1/10). What changed, and what I pushed back on: · 已处理最新评审反馈(第 1/10 轮)。改动内容与我反驳保留之处如下:

Autofix round — PR #13481 (issue #13479)

Commit 6aebce3d85 — fix(release): harden the docker disk reclaim and data-root floor gate (#13479). All 8 round-1 findings (every one a Suggestion from the automated reviewer) were verified against the code and implemented; nothing was declined, deferred, or escalated. Changes stayed inside the PR's footprint: .github/scripts/run-release-docker-integration.sh and scripts/tests/release-workflow.test.js.

Dispositions

  • R1-7 (rc:4190525808) — dangling images unreclaimed → implemented. Added the sweep-shaped docker image prune --force --filter 'until=24h' (warn-only, same shape as ecs-runner/qwen-docker-cleanup.sh:44) between the labelled prune and the builder prune, with a comment on why the labelled prune cannot reach that class. --no-prune on the build call was left alone, as the finding required.
  • R1-2 (rc:4190525817) — unbounded prune inside the host build mutex → implemented. The builder prune is now timeout 20m docker builder prune --all --force --filter 'until=24h', keeping the || echo "::warning::…" tail. The comment now states the deliberate omission of --keep-storage (a reserve, not a quota — it would reclaim nothing on exactly the below-30 GB hosts this line exists for) and why the bound matters (the E2E lane waits only 30 minutes on the same mutex; run 33637097713).
  • R1-13 (rc:4190525834) — only unscoped check-disk-floor.sh invocation repo-wide → implemented (scoped). The gate now runs self-hosted only, matching all 16 other call sites and the helper's own header, with the divergence stated in the function comment (an ephemeral hosted runner starts with an order of magnitude more free disk than the 8 GiB floor — the finding's own measurement was ~82 GiB on ubuntu-latest). Judgment call: the finding offered scope-or-document; I chose scoping because it matches the repo convention and the helper's documented purpose, and the failure mode this gate exists for (a saturated persistent pool host) does not exist on an ephemeral hosted VM. The hosted lane's behavior is now pinned by a test asserting no gate, no docker info, and no flock calls.
  • R1-3 (rc:4190525842) — three silent gate-skip paths → implemented. The gate logic moved into check_docker_data_root_floor(), which warns (exit 0, build proceeds) when the helper is absent at the released ref and when the docker data root cannot be resolved or is not a readable directory. The pre-existing [ -f ] relative-to-step-cwd convention was kept, as the finding instructed.
  • R1-5 (rc:4190525845) — gate samples only before the build → implemented (re-gate). The same gate runs a second time after npm run build:sandbox and before the two vitest phases, with a comment recording that run 37374675168 died in the vitest phase ~15 minutes after the build. Narrow variant: the image-present path keeps its no-gate contract, which the finding presented as the default.
  • R1-8 (rc:4190525852) — stub answered any info invocation → implemented. The docker stub now answers only the exact info --format '{{.DockerRootDir}}' contract (anything else prints nothing and exits 1), and the self-hosted execution test pins expect(calls).toContain('docker info --format {{.DockerRootDir}}'). A template typo now drops the gate lines and reds the suite (probe P8 below). DISK_FLOOR_MIN_FREE_KB was not renamed.
  • R1-9 (rc:4190525857) — prune failure path unpinned → implemented. New pruneFails harness option makes the stub fail builder prune only; the case asserts exit 0, the ::warning::docker build cache cleanup failed line, and that the build and both vitest runs still happen.
  • R1-10 (rc:4190525862) — self-hosted lock path never executed under test → implemented. New runnerEnvironment option (default github-hosted) plus flock and timeout stubs (both absent on stock macOS, which runs this suite on merge_group). The self-hosted case asserts the full recorded order — daemon lock → coordinator lock → build mutex → labelled prune → dangling prune → cache prune → docker info → pre-build gate → build → unlock 7 → post-build gate → unlock 8 → vitest — and the gate-trip case now asserts a non-zero exit with no unlock lines past the trip (the mutant-C3 shape the finding measured as uncovered). Fd 9 is correctly never asserted unlocked.

Verification

  • bash -n .github/scripts/run-release-docker-integration.sh — passed.
  • npx prettier --check scripts/tests/release-workflow.test.js — passed (the .sh has no prettier parser, as before).
  • npm run build — passed (exit 0).
  • npm run typecheck — passed (exit 0).
  • npm run lint — passed (exit 0).
  • shellcheck v0.11.0 (pinned version, run directly on the script; the repo's node scripts/lint.js --shellcheck lane cannot run in this sandbox because file(1) is absent) — clean.
  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/release-workflow.test.js — 76 passed (baseline pre-round: 71 passed; the 3 pre-existing execution tests keep their names, 5 cases added). Re-run after the commit (which includes the pre-commit hook's prettier/eslint pass): still 76 passed, working tree clean.
  • Mutation probes (each: mutate the script, re-run the suite, confirm red, restore; suite green again after restore):
    • P1 dangling image prune removed → 3 failed (caught).
    • P2 timeout 20m bound removed → 1 failed (caught).
    • P3 self-hosted scoping removed → 1 failed (caught).
    • P4 helper-absent warning removed → 1 failed (caught).
    • P5 unreadable-root warning removed → 2 failed (caught).
    • P6 post-build re-gate removed → 2 failed (caught).
    • P7 prune || echo "::warning::…" tail removed → 1 failed (caught).
    • P8 {{.DockerRootDir}} → {{.DockerRoot}} typo → 2 failed (caught).
    • P9 early exit before flock --unlock 7 (the finding's mutant C3) → 4 failed (caught).
  • Pre-round check (this round's tests against the pre-round script): 6 of the changed/new tests fail pre-round — they pin the new behavior; the 3 that stay green (fails before…, skips the reclaim…, continues with a warning when the BuildKit cache prune fails) pin pre-existing behavior this round did not alter.
  • Not run: integration tests after npm run bundle — the touched behavior (a release-workflow shell script and its harness) is not exercised through the bundled CLI. No settings source changed, so npm run generate:settings-schema was not needed.
中文说明

Autofix 本轮处理 — PR #13481(issue #13479)

提交 6aebce3d85 —— fix(release): harden the docker disk reclaim and data-root floor gate (#13479)。第 1 轮评审的全部 8 条发现(均为自动评审的 Suggestion)都已对照代码核实并实现;没有驳回、延期或升级给维护者的条目。改动严格留在本 PR 的足迹内:.github/scripts/run-release-docker-integration.sh 与 scripts/tests/release-workflow.test.js。

处理结论

  • R1-7(rc:4190525808)——悬空镜像无法回收 → 已实现。 在带标签 prune 与 builder prune 之间加入与 sweep 同形的 docker image prune --force --filter 'until=24h'(只告警不失败,形状同 ecs-runner/qwen-docker-cleanup.sh:44),并用注释说明为何带标签的 prune 触达不到这一类。按该发现的要求,构建调用上的 --no-prune 保持不变。
  • R1-2(rc:4190525817)——主机互斥锁内的 prune 无时间上限 → 已实现。 builder prune 现为 timeout 20m docker builder prune --all --force --filter 'until=24h',并保留 || echo "::warning::…" 兜底。注释现在写明有意省略 --keep-storage(它是保留额度而非配额——在缓存不足 30 GB 的主机上,也就是本行为之存在的主机上,它什么也不会回收),以及加上限的原因(E2E 通道在同一把互斥锁上只等 30 分钟;run 33637097713)。
  • R1-13(rc:4190525834)——全仓库唯一未限定作用域的 check-disk-floor.sh 调用 → 已实现(限定为 self-hosted)。 该门现在仅在 self-hosted 上运行,与其余 16 处调用点及辅助脚本自身头部注释一致,并在函数注释中说明这一分歧(临时 hosted runner 初始可用磁盘比 8 GiB 下限高一个数量级——该发现自己的实测为 ubuntu-latest 上约 82 GiB)。判断说明:该发现给出了「限定作用域或写明文档」两个选项;我选择限定作用域,因为它符合仓库惯例与该辅助脚本文档的用途,且该门针对的失败模式(饱和的常驻机器池主机)在临时 hosted 虚拟机上不存在。hosted 通道的行为现由一条测试固化:断言不出现 gate、docker info 与 flock 调用。
  • R1-3(rc:4190525842)——三条静默的门跳过路径 → 已实现。 门逻辑移入 check_docker_data_root_floor():当辅助脚本在被发布 ref 上不存在、以及 docker 数据根无法解析或不是可读目录时,都会告警(退出码 0,构建继续)。按该发现的指示,保留了既有的 [ -f ] 相对步骤 cwd 解析的约定。
  • R1-5(rc:4190525845)——门只在构建前采样一次 → 已实现(构建后再次设门)。 同一道门在 npm run build:sandbox 之后、两个 vitest 阶段之前再执行一次,并以注释记录 run 37374675168 实际死于构建结束约 15 分钟后的 vitest 阶段。采用窄化变体:镜像已存在路径保持不设门的约定,这也是该发现列出的默认方案。
  • R1-8(rc:4190525852)——桩对任何 info 调用都返回装置根目录 → 已实现。 docker 桩现在只响应精确的 info --format '{{.DockerRootDir}}' 契约(其他任何调用不输出并退出 1),且 self-hosted 执行测试固化了 expect(calls).toContain('docker info --format {{.DockerRootDir}}')。模板拼写错误现在会让门消失并使套件变红(见下方探针 P8)。DISK_FLOOR_MIN_FREE_KB 未改名。
  • R1-9(rc:4190525857)——prune 失败路径未被固化 → 已实现。 新增 pruneFails 装置选项,让桩仅在 builder prune 上失败;用例断言退出码 0、出现 ::warning::docker build cache cleanup failed 告警,且构建与两次 vitest 仍照常执行。
  • R1-10(rc:4190525862)——self-hosted 锁路径从未在测试下执行 → 已实现。 新增 runnerEnvironment 选项(默认 github-hosted)以及 flock 与 timeout 桩(原生 macOS 两者皆无,而 merge_group 会在 macOS 上运行本套件)。self-hosted 用例断言完整记录顺序——daemon 锁 → 协调锁 → 构建互斥锁 → 带标签 prune → 悬空 prune → 缓存 prune → docker info → 构建前门 → 构建 → 解锁 7 → 构建后门 → 解锁 8 → vitest——门触发的用例现在还断言非零退出且触发点之后没有任何解锁记录(即该发现实测未覆盖的 mutant-C3 形态)。fd 9 按设计要求不会被断言解锁。

验证

  • bash -n .github/scripts/run-release-docker-integration.sh —— 通过。
  • npx prettier --check scripts/tests/release-workflow.test.js —— 通过(.sh 一如既往没有 prettier 解析器)。
  • npm run build —— 通过(退出码 0)。
  • npm run typecheck —— 通过(退出码 0)。
  • npm run lint —— 通过(退出码 0)。
  • shellcheck v0.11.0(仓库锁定版本,直接对脚本运行;仓库的 node scripts/lint.js --shellcheck 通道因本沙箱缺少 file(1) 而无法运行)—— 无告警。
  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/release-workflow.test.js —— 76 项通过(本轮前基线:71 项通过;3 个既有执行测试保留原名,新增 5 个用例)。提交后(含 pre-commit 钩子的 prettier/eslint 处理)重跑:仍为 76 项通过,工作区干净。
  • 变异探针(每个:改动脚本、重跑套件、确认变红、还原;还原后套件复绿):
    • P1 删除悬空镜像 prune → 3 项失败(已捕获)。
    • P2 删除 timeout 20m 上限 → 1 项失败(已捕获)。
    • P3 删除 self-hosted 限定 → 1 项失败(已捕获)。
    • P4 删除辅助脚本缺失告警 → 1 项失败(已捕获)。
    • P5 删除数据根不可读告警 → 2 项失败(已捕获)。
    • P6 删除构建后再次设门 → 2 项失败(已捕获)。
    • P7 删除 prune 的 || echo "::warning::…" 兜底 → 1 项失败(已捕获)。
    • P8 {{.DockerRootDir}} → {{.DockerRoot}} 拼写错误 → 2 项失败(已捕获)。
    • P9 在 flock --unlock 7 之前提前退出(该发现的 mutant C3)→ 4 项失败(已捕获)。
  • 轮前检查(用本轮测试跑轮前脚本):6 个改动/新增测试在轮前失败——它们固化的正是新行为;保持绿色的 3 个(fails before…、skips the reclaim…、continues with a warning when the BuildKit cache prune fails)固化的是本轮未改变的既有行为。
  • 未运行:npm run bundle 后的集成测试——所改行为(发布工作流 shell 脚本及其测试装置)不经过打包后的 CLI。未改动 settings 源文件,故无需运行 npm run generate:settings-schema。

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。

🧵 Resolved all 8 selected review thread(s). · 已关闭全部选中的 8 条评审线程。

Re-review when you have a moment. After round 10 this bot stops and leaves the PR for a human. · 有空请复审;第 10 轮后本 bot 停止并将 PR 交给人工。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.25.0

@yiliang114 yiliang114 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 2 verified against 6aebce3: all eight round-1 suggestions are addressed (timeout on the builder prune, dangling-image prune, self-hosted scoping, warned skip paths, the post-build re-gate, the exact-match docker info stub, the prune-failure test, and self-hosted execution tests). I also ran the suite locally on macOS — 81 passed / 1 skipped — which closes the "never ran on macOS" gap disclosed last round. ShellCheck clean; CI green on this head.

Two suggestions inline. Neither is a defect in the incident fix itself — the reclaim + gate direction looks correct — both are hardening gaps measured against the PR's own stated rationale. I'm not approving with open findings; happy to approve once you've taken a look, whether that means fixing or declining with a reason.

Comment thread .github/scripts/run-release-docker-integration.sh Outdated
Comment thread .github/scripts/run-release-docker-integration.sh Outdated
@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🤖 Addressed the latest review feedback (round 2/10). What changed, and what I pushed back on: · 已处理最新评审反馈(第 2/10 轮)。改动内容与我反驳保留之处如下:

Round summary

Both inline findings from @yiliang114's round-2 review (CHANGES_REQUESTED) are implemented in 32e8b5a. No base-branch conflict (--conflict false; no merge performed). Diff this round: 2 files, +87/−22 — inside the window's growth budget.

rc:4191354944 — unbounded daemon round-trips under the host build mutex (P3) — implemented

The two docker image prune calls and the docker info inside check_docker_data_root_floor ran while fd 7 (the exclusive host build mutex) was held, carrying no time bound; on a wedged daemon either could hold the mutex past the E2E lane's 30-minute flock --wait 1800. Applied the suggested bounds:

  • timeout 20m docker image prune --all --force … (labelled prune — the reviewer noted this pre-existing call has the same exposure, so both prunes are bounded for consistency)
  • timeout 20m docker image prune --force … (dangling prune)
  • timeout 60 docker info --format '{{.DockerRootDir}}' inside the gate

A timeout degrades exactly like any other failure of that call: the prunes fall into their existing || echo "::warning::…" continuation, and a docker info timeout yields an empty root, which the gate already handles as warn-and-skip. The builder-prune comment's rationale sentence was generalized to cover every daemon call in the branch instead of only that prune.

Witness: the static pin test now requires the timeout 20m / timeout 60 prefixes on all three lines. Mutation probe: stripping the three prefixes fails that test; restored → green.

rc:4191354951 — gates skipped on the cached-image re-run path (P3) — implemented

Both gate calls lived inside the image-missing branch, so a re-run landing on the same still-saturated host — where the image is now cached — went straight to the vitest phase, the actual #13479 death site, ungated. The post-build gate is hoisted out of the branch to directly before the two vitest runs, so every path (cold build, cached re-run) passes the data-root gate before vitest writes container layers to that filesystem. The pre-build gate stays inside the branch: with no build there is nothing to pre-gate. The reviewer is right that this widens coverage past the PR body's original scoping, but the move is two lines, closes the re-run edge of the very incident this PR fixes, and keeps the gate count at two on the cold path.

Witness: two new stubbed-execution tests on the cached + self-hosted path — one asserts the gate runs exactly once before the first vitest invocation while the reclaim and build are skipped; one asserts a gate trip blocks vitest entirely. Mutation probe: deleting the hoisted gate call fails both new tests plus the static gate-placement pin and the cold-path ordering test; restored → green.

Verification

  • bash -n .github/scripts/run-release-docker-integration.sh — passed
  • npx prettier --check scripts/tests/release-workflow.test.js .github/scripts/run-release-docker-integration.sh — passed
  • COREPACK_HOME=/tmp/corepack-home npx vitest run --config ./scripts/tests/vitest.config.ts release-workflow.test — 84 passed (78 in release-workflow.test.js, including the 2 new cached-path tests; 6 in cua-driver-release-workflow.test.js). COREPACK_HOME is pointed at a writable tmp dir because this sandbox denies writes to the default corepack cache; the suite itself is the repo's standard test:scripts config.
  • Mutation probe A (strip the three new timeout prefixes from the script) — the static pin test failed as designed; restored, suite green
  • Mutation probe B (delete the hoisted gate call) — 4 tests failed as designed (static pin, cold-path ordering, both new cached-path tests); restored, suite green
  • npm run build — passed
  • npm run typecheck — passed
  • npm run lint — passed

Not run: ShellCheck (no binary on this runner; syntax is covered by bash -n and the stubbed execution tests — CI runs ShellCheck). The real self-hosted docker lane cannot run off the pool by definition; CI remains the final gate.

中文说明

本轮摘要

@yiliang114 第二轮评审(CHANGES_REQUESTED)中的两条行内意见均已在 32e8b5a 中实现。与 base 分支无冲突(--conflict false,未做合并)。本轮 diff:2 个文件,+87/−22,在窗口增长预算之内。

rc:4191354944 —— 宿主机构建互斥锁内未设时限的 daemon 往返调用(P3)—— 已实现

两个 docker image prune 调用以及 check_docker_data_root_floor 内的 docker info 都在持有 fd 7(宿主机排他构建互斥锁)期间执行且没有时间上限;一旦 daemon 卡死,任一调用都可能把互斥锁拖到超过 E2E 通道 30 分钟的 flock --wait 1800。已按建议加上时限:

  • timeout 20m docker image prune --all --force …(带标签的清理 —— 评审者指出这个原有调用有同样的暴露面,因此两个 prune 一并加上时限以保持一致)
  • timeout 20m docker image prune --force …(悬空镜像清理)
  • 门禁函数内的 timeout 60 docker info --format '{{.DockerRootDir}}'

超时后的降级路径与该调用任何其他失败完全一致:prune 落入已有的 || echo "::warning::…" 继续执行分支;docker info 超时会产生空 root,门禁本就走 warn-and-skip。builder prune 注释中的理由句已推广为覆盖该分支内的所有 daemon 调用,而不仅是那一条 prune。

佐证:静态 pin 测试现在要求三行都带 timeout 20m / timeout 60 前缀。变异探针:去掉这三个前缀会让该测试如期失败;还原后恢复全绿。

rc:4191354951 —— 镜像已缓存的重跑路径跳过门禁(P3)—— 已实现

两个门禁调用原先都在 image-missing 分支内,因此当重跑落在同一台仍然磁盘饱和的主机上、镜像已被缓存时,会直接跳到 vitest 阶段 —— 也就是 #13479 真正的死亡现场 —— 且完全无门禁。现将构建后门禁提升到分支之外、紧贴两个 vitest 运行之前,使每条路径(冷构建、命中缓存的重跑)在 vitest 向该文件系统写入容器层之前都先过数据根目录门禁。构建前门禁保留在分支内:没有构建就没有需要前置门禁的对象。评审者说得对,这确实超出了 PR 描述最初的范围界定,但这次移动只有两行,闭合的正是本 PR 所修复事故的重跑边角,且冷路径上的门禁总数仍是两次。

佐证:新增两个基于 stub 的执行测试覆盖 缓存命中 + self-hosted 路径 —— 一个断言在 reclaim 与 build 均被跳过的前提下,门禁恰好在首个 vitest 调用之前运行一次;另一个断言门禁触发失败会完全阻断 vitest。变异探针:删除提升后的门禁调用会让两个新测试、静态位置 pin 以及冷路径顺序测试共 4 个测试如期失败;还原后恢复全绿。

验证

  • bash -n .github/scripts/run-release-docker-integration.sh —— 通过
  • npx prettier --check scripts/tests/release-workflow.test.js .github/scripts/run-release-docker-integration.sh —— 通过
  • COREPACK_HOME=/tmp/corepack-home npx vitest run --config ./scripts/tests/vitest.config.ts release-workflow.test —— 84 个通过(release-workflow.test.js 78 个,含 2 个新增缓存路径测试;cua-driver-release-workflow.test.js 6 个)。COREPACK_HOME 指向可写临时目录,因为本沙箱禁止写默认 corepack 缓存;套件本身仍是仓库标准的 test:scripts 配置。
  • 变异探针 A(去掉脚本中三处新增 timeout 前缀)—— 静态 pin 测试如期失败;还原后套件全绿
  • 变异探针 B(删除提升后的门禁调用)—— 4 个测试如期失败(静态 pin、冷路径顺序、两个新增缓存路径测试);还原后套件全绿
  • npm run build —— 通过
  • npm run typecheck —— 通过
  • npm run lint —— 通过

未运行:ShellCheck(本 runner 无此二进制;语法由 bash -n 与 stub 执行测试覆盖 —— CI 会跑 ShellCheck)。真实的 self-hosted docker 通道离开资源池本质上无法运行;CI 仍是最终门禁。

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。

🧵 Resolved all 2 selected review thread(s). · 已关闭全部选中的 2 条评审线程。

Re-review when you have a moment. After round 10 this bot stops and leaves the PR for a human. · 有空请复审;第 10 轮后本 bot 停止并将 PR 交给人工。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.25.0

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🤖 Reviewed the latest feedback — no changes needed. Why, point by point: · 已审阅最新反馈——无需改动。逐点说明原因如下:

🤖 No action this round (round 3/10) — the round was dispatched by two failed checks in the review automation itself, not by anything in this PR. · 本轮无改动(第 3/10 轮)—— 触发本轮的是评审自动化自身的两个失败检查,而非本 PR 的内容。

Round summary

No new review feedback exists since the last evaluation (2026-10-06T03:50:48Z): no reviews, no inline comments, no issue-level comments. Both round-2 inline findings (rc:4191354944, rc:4191354951) were implemented in 32e8b5a and their threads were resolved last round. The only listed items are two failed checks, both in the 🧐 Qwen Pull Request Review workflow (the review automation pipeline, run 37428041959):

  • review-pr — FAILURE after 28 seconds (07:18:12Z → 07:18:40Z)
  • fallback-comment — FAILURE after 34 seconds (07:18:43Z → 07:19:17Z)

Why these failures are not actionable from this PR

The failing jobs never execute anything this PR changes. The PR touches exactly two files:

  • .github/scripts/run-release-docker-integration.sh — invoked only by release.yml (lines 572/578); the review workflow never runs it.
  • scripts/tests/release-workflow.test.js — run by the CI Test job, which is green on this head (see below).

The timing rules out a review of the diff. review-pr failed 28 seconds after start — long before checkout, dependency install, CLI build, or the review step itself could run. Its first steps are self-hosted runner health probes (directory-writability checks, workspace-ownership restore, stale-state cleanup) on the same shared self-hosted pool whose disk saturation is the subject of issue #13479. A fast failure there is an infrastructure event, not a verdict on this diff. The same workflow reviewed this PR successfully twice before on larger diffs.

fallback-comment runs on GitHub-hosted ubuntu-latest and exists only to post a comment when review-pr fails. Its only failure paths are GitHub API calls (the dedup lookup fails closed after three retries rather than risking a duplicate comment). A 34-second failure there is an API/auth transient in the automation, again unrelated to this PR's content.

Every check that does execute this PR's content is green on 32e8b5a: Lint & Static ✅, Test (ubuntu-latest, Node 22.x) ✅, Integration Tests (no-AK, No Sandbox) ✅, Desktop Shell (ubuntu-22.04) ✅, Desktop Shell (windows-2022) ✅, web-shell E2E Smoke ✅.

The boundary applies. The review-automation workflow (.github/workflows/qwen-code-pr-review.yml) is CI/verification machinery this PR is not about; autofix is not permitted to modify it, and no change to the release script or its tests could affect these two jobs anyway.

What resolves this

A maintainer re-run of the failed jobs on run 37428041959 (Actions → 🧐 Qwen Pull Request Review → Re-run failed jobs). That needs no code change; if the failures persist across re-runs, they belong to the review-automation infrastructure (runner pool health / API credentials), tracked outside this PR.

Verification

No code changed this round, so no build/test re-run was triggered; the working tree is clean and byte-identical to head 32e8b5a, the exact commit the green checks above ran against. Evidence gathered this round:

  • git status --porcelain=v1 --untracked-files=all — clean; HEAD = 32e8b5a
  • git diff origin/main...HEAD — re-verified both round-2 findings are implemented in the head commit (timeout bounds on the two docker image prune calls and the gate's docker info; post-build data-root gate hoisted ahead of the vitest runs so the cached-image re-run path is gated)
  • Check-run metadata for run 37428041959 — per-job start/completion timestamps quoted above
  • grep over .github/workflows/ — confirmed the touched script is referenced only by release.yml, never by the review workflow
中文说明

本轮摘要

自上次评估(2026-10-06T03:50:48Z)以来没有新的评审反馈:没有新的 review、没有行内评论、没有 issue 级评论。第二轮的两条行内意见(rc:4191354944、rc:4191354951)已在 32e8b5a 中实现,其线程已在上一轮关闭。本次列出的唯一条目是两个失败的检查,且都来自 🧐 Qwen Pull Request Review 工作流(评审自动化管线,运行 37428041959):

  • review-pr —— 启动 28 秒后失败(07:18:12Z → 07:18:40Z)
  • fallback-comment —— 启动 34 秒后失败(07:18:43Z → 07:19:17Z)

为什么这些失败无法从本 PR 侧处理

失败的 job 从不执行本 PR 改动的任何内容。 本 PR 只触及两个文件:

  • .github/scripts/run-release-docker-integration.sh —— 仅被 release.yml(第 572/578 行)调用;评审工作流从不运行它。
  • scripts/tests/release-workflow.test.js —— 由 CI 的 Test job 运行,该 job 在当前 head 上是绿的(见下)。

耗时上也不可能是对 diff 的评审结论。 review-pr 在启动 28 秒后即失败 —— 远早于 checkout、依赖安装、CLI 构建或评审步骤本身可能完成的时间。该 job 的最前面几步是 self-hosted runner 健康探针(目录可写性检查、工作区属主修复、残留状态清理),运行在同一个共享 self-hosted 资源池上 —— 而该池的磁盘饱和正是 issue #13479 的主题。在这个阶段的快速失败属于基础设施事件,而非对本 diff 的评审结论。同一工作流此前已在更大的 diff 上成功评审过本 PR 两次。

fallback-comment 运行在 GitHub 托管的 ubuntu-latest 上,其唯一职责是在 review-pr 失败时发布一条评论。它仅有的失败路径是 GitHub API 调用(去重查询在三次重试后选择失败关闭,以避免重复评论)。34 秒的失败属于自动化链路中的 API/凭据瞬时故障,同样与本 PR 内容无关。

所有真正执行本 PR 内容的检查在 32e8b5a 上均为绿: Lint & Static ✅、Test (ubuntu-latest, Node 22.x) ✅、Integration Tests (no-AK, No Sandbox) ✅、Desktop Shell (ubuntu-22.04) ✅、Desktop Shell (windows-2022) ✅、web-shell E2E Smoke ✅。

边界规则适用。 评审自动化工作流(.github/workflows/qwen-code-pr-review.yml)属于本 PR 并未涉及的 CI/验证机制;autofix 无权修改它,而且对发布脚本或其测试的任何改动本来也不可能影响这两个 job。

如何解决

由维护者在运行 37428041959 上重跑失败的 job(Actions → 🧐 Qwen Pull Request Review → Re-run failed jobs)。这不需要任何代码改动;如果重跑后仍持续失败,则属于评审自动化基础设施问题(runner 池健康 / API 凭据),应在本 PR 之外跟踪处理。

验证

本轮无代码改动,因此未触发构建/测试重跑;工作树干净,与 head 32e8b5a 逐字节一致 —— 正是上述绿色检查所运行的那个提交。本轮收集的证据:

  • git status --porcelain=v1 --untracked-files=all —— 干净;HEAD = 32e8b5a
  • git diff origin/main...HEAD —— 复核确认第二轮两条意见均已实现在 head 提交中(两个 docker image prune 调用与门禁内 docker info 的 timeout 上限;构建后数据根目录门禁已提升到 vitest 运行之前,使命中缓存的重跑路径同样被门禁覆盖)
  • 运行 37428041959 的检查元数据 —— 上文引用的各 job 起止时间
  • 对 .github/workflows/ 的 grep —— 确认被改脚本仅被 release.yml 引用,评审工作流从不引用它

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.25.0

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🔀 Base updated: red check(s) [review-pr] pass on current main — merged current main via update-branch; CI will re-run.

中文说明

🔀 已更新 base:红色检查 [review-pr] 在当前 main 上通过 —— 已通过 update-branch 合入当前 main,CI 将重新运行。

@qwen-code-review-bot

Copy link
Copy Markdown
Collaborator

Qwen Code review did not complete successfully. The review pipeline failed before a review could be posted. A transient error is retried automatically; if you are seeing this, retry with @qwen-code /review. See workflow logs.

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🤖 Reviewed the latest feedback — no changes needed. Why, point by point: · 已审阅最新反馈——无需改动。逐点说明原因如下:

Autofix review round — no action

No code changes this round. The feedback window contained no actionable findings on the PR's code:

  • Inline comments: none.
  • Reviews: none.
  • Issue-level comments: one — [ic:6023889269], the automated reviewer's fallback notice that the review-pr pipeline itself failed before a review could be posted (<!\-\- qwen-review-fallback -->). This is infrastructure state, not a finding on the diff: the notice says a transient error is retried automatically, and re-triggering the review (@qwen-code /review) is a GitHub-side action this round has no credentials to perform. The next evaluation cycle re-runs the review; any findings it produces will be addressed then.
  • Failed checks: review-pr: FAILURE — the same review-pipeline failure above, not a build, lint, typecheck, or test failure on this branch.
  • Diff growth this window: source 0 / test 0 — nothing added.

Verification

Since no files changed, the full build/typecheck/lint battery was not re-run. The PR's own focused test suite was executed against the current branch head to confirm the branch is healthy:

  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/release-workflow.test.js — 78/78 passed (run with COREPACK_HOME=/tmp/corepack-cache because this sandbox cannot write the default corepack cache under /home/github-runner/.cache; the warmup is environment setup, not a test result).

The branch remains exactly as the last evaluation left it, ready for the review pipeline's retry.

中文说明

Autofix 审查轮次 —— 无需操作

本轮没有代码改动。反馈窗口中没有任何针对本 PR 代码的可执行发现:

  • 行内评论: 无。
  • Reviews: 无。
  • Issue 级评论: 一条 —— [ic:6023889269],自动审查器的回退通知,说明 review-pr 流水线本身在发布审查结果之前失败了(<!\-\- qwen-review-fallback -->)。这是基础设施状态,而不是针对 diff 的发现:该通知说明瞬时错误会自动重试,而重新触发审查(@qwen-code /review)是 GitHub 侧的操作,本轮没有相应凭据可以执行。下一个评估周期会重新运行审查;届时产生的任何发现都会被处理。
  • 失败的检查: review-pr: FAILURE —— 即上述同一个审查流水线失败,并非本分支上的构建、lint、类型检查或测试失败。
  • 本窗口 diff 增长: 源码 0 行 / 测试 0 行 —— 没有新增内容。

验证

由于没有文件改动,未重新运行完整的 build/typecheck/lint 流程。针对当前分支顶端运行了 PR 自身的聚焦测试套件,以确认分支状态健康:

  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/release-workflow.test.js —— 78/78 通过(运行时设置了 COREPACK_HOME=/tmp/corepack-cache,因为本沙箱无法写入 /home/github-runner/.cache 下默认的 corepack 缓存目录;该 warmup 属于环境准备,不是测试结果)。

分支保持在上次评估时的原状,等待审查流水线重试。

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.25.0

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🔀 Base updated: red check(s) [review-pr] pass on current main — merged current main via update-branch; CI will re-run.

中文说明

🔀 已更新 base:红色检查 [review-pr] 在当前 main 上通过 —— 已通过 update-branch 合入当前 main,CI 将重新运行。

@wenshao

wenshao commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator

Verdict: merge-ready — 69/69 scripted assertions passed, 0 unexpected failures. Verified head 2560f56274fe27523abadf0a523d8ba6c74c9fbd.

中文摘要

结论:merge-ready(69/69 通过)。本轮在真实环境(Linux aarch64,Node 24)对 PR 的中央主张做了 A/B 证明:

  • 套件级 A/B:head 上 release-workflow 套件 83 绿 1 跳过,11 个新测试全部执行且通过;把 head 测试文件原样放到基线树(f3385785)上重放,恰好 10 个新机制测试变红(第 11 个"镜像已存在则跳过回收与门"在基线上也绿——它钉的是未变行为),失败形态为预期断言红(如 expect(result.status).not.toBe(0)),非收集错误。
  • 真实 helper 线束:PR 自带测试把 check-disk-floor.sh 换成了 stub;本轮用真实脚本 + 真实 helper 组合驱动(仅 docker/git/npm/npx/flock 为 stub)。head:8 GiB 下限触发时 exit 1、构建与 vitest 均被阻断、输出 ::error::Disk floor breached(单元 B/H);下限默认 8388608 与目录实读均正确(A/I);DISK_FLOOR_MIN_FREE_KB 覆盖生效(B vs C);helper 的 inode 腿同样能触发门(D);docker info 返回垃圾时降级为 warn-skip(G);hosted runner 上门被正确关闭(F)。基线对照 E:同样的巨大下限环境下基线脚本没有任何门,构建照常——证明机制是 PR 引入的。
  • 变异矩阵:4/4 全部被杀(删 cache prune / 门降级为 warn-only / 删 pre-vitest 门 / 下限减半),M4 为同文件阳性对照。
  • 附带 corroboration:事故 run 37374675168 的 "Integration Tests (Docker)" 确于 2026-10-05 失败(21:18→21:46);本机真实 docker daemon 对 {{.DockerRootDir}} 模板返回 /var/lib/docker,真实 helper 对其输出正确 DISKFLOOR 样本;24h 过滤器与 qwen-docker-cleanup.sh(含 30 GB keep-storage)及 e2e.yml 既有行一致;git merge-tree 与当前 main 干净合并,且 main 自 merge-base 起未触碰这两个文件。

发现(均不阻塞):① PR 描述称"镜像已被并发构建时行为不变",但 pre-vitest 门刻意位于 image-missing 分支之外(代码注释与两个新测试均如此钉),缓存镜像路径现在会在 vitest 前多一次门检查——代码与测试是对的,仅描述措辞不准,建议改正文一句话;② 降级路径(helper 缺失/docker info 失败/根目录不可读)按设计 warn-skip;③ 门同时带 helper 默认的 inode 腿(100000),PR 只覆盖了 KB 下限,属一致的额外覆盖。

未覆盖:发布 Docker 通道端到端真实构建(未执行 docker 构建;8 GiB 的尺寸判断接受作者结论,可 env 覆盖);固定 ShellCheck 0.11.0 门(lint.js 无 linux/aarch64 平台项,本机无法跑固定版本;用发行版 0.9.0 旁证:默认规则集两臂均干净,CI 精确旗标两臂同样失败于既有风格提示,A/A 证明与 PR 无关);scripts 全套件(改动面仅 release-workflow.test.js,本轮跑的是该定向套件);事故 run 的步骤级日志(annotations API 404,已用 job 级事实 corroborate)。

Evidence: 01-ab-suite-head-vs-base.png · 02-mutation-matrix.png · 03-real-helper-harness.png (appended below)

Central claim and A/B

Claim under test: on the self-hosted release Docker lane, an image-missing run now prunes dangling images and BuildKit cache (both until=24h, timeout 20m-bounded) and gates the Docker data root at 8 GiB before npm run build:sandbox; a tripped gate fails fast before the build; a second gate guards the vitest phase (the actual #13479 death site) on every self-hosted run, including cached-image re-runs.

Cell Oracle Base f3385785 Head 2560f56274
release-workflow suite (head tests on both trees) exit / red names 10 red, exactly the new-mechanism tests; failure shape = intended assertions (e.g. expect(result.status).not.toBe(0)), 73 green 83 pass / 1 skip, all 11 new tests executed (verbose ✓ census)
Real-helper harness, image-missing, floor breached (B vs E) build ran? / DISKFLOOR / exit / ::error:: build ran, 0 samples, exit 0 — nothing gates build skipped, 1 sample floor_kb[8388608], exit 1, ::error::Disk floor breached, no lock unlocks
Cached-image re-run, floor breached (H) vitest runs? (base has no gate at all) vitest blocked, exit 1, 1 sample
Happy path, self-hosted, image missing (A) order + counts — image prune → dangling prune → builder prune → docker info → gate → build → gate → vitest×2, exit 0
Override honored (C) floor_kb value — floor_kb[1], build ran
Inode leg (D) gate trips on inodes — exit 1, build skipped
Hosted runner, floor env huge (F) gate scoped off — 0 DISKFLOOR samples, build ran
docker info garbage (G) degrade path — ::warning::…gate skipped, build ran

The PR's own suite stubs check-disk-floor.sh; the harness above composes the real script with the real helper (only docker/git/node/npm/npx/timeout/flock stubbed), so the env-var name, argv shape, exit-code propagation through set -euo pipefail, and the ::error:: annotation are pinned by execution. 58/58 assertions pass (41 head A–G + 13 head H–I + 4 base E).

Mutation matrix (vacuity proof, head)

Mutation (in .github/scripts/run-release-docker-integration.sh) Result Red tests
M1 delete the BuildKit cache prune killed 4: text pin, hosted reclaim, build order, prune-failure warning
M2 degrade gate to warn-only (` echo …`)
M3 delete the pre-vitest gate call killed 4: text pin (gate count), build order, both cached-image gate tests
M4 floor 8388608 → 4194304 (positive control) killed 3: text pin + both floor-value assertions

Script restored and sha256-verified after every row. Zero survivors, so no adjudication needed; M4 proves the runner collects tests that pin this file.

Corrections

  • PR body, "Risk & Scope"/overview: "Both behaviors are scoped to the branch that actually builds the image; when a concurrent run already built the image for the same revision, nothing changes" — and the Reviewer Test Plan's "an already-present image skips both". The pre-vitest check_docker_data_root_floor call deliberately sits outside the image-missing branch (the code comment says so, and the tests gates the cached-image re-run path… / blocks the vitest phase… pin it). On the cached-image path something does change: one data-root gate before vitest. Verified live in cells H/I with the real helper. The code and tests are correct and intentional — this is a description nit only; a one-line body edit would fix it.

Findings

None blocking. Observations, ordered:

  1. Degrade paths are warn-and-skip by design (missing helper on old release refs, failed/timeout 60 docker info, non-directory root) — cell G and the PR's three warn-path tests pass. On a host with a wedged daemon the gate never fires, but that host's build fails anyway; consistent with check-disk-floor.sh's existing "letting the job proceed" philosophy. No action.
  2. The gate is two-legged: the helper's inode floor (default 100000) also applies to the docker root; only the KB floor is overridden to 8 GiB. Cell D proved the inode leg trips the gate. Coherent bonus coverage — noted so reviewers know the gate is not space-only.
  3. Consistency claims verified: the new prunes match the daily sweep qwen-docker-cleanup.sh (until=24h; the sweep additionally carries --keep-storage 30GB, and the PR's comment correctly explains omitting it — a reserve reclaims nothing on the saturated hosts this line exists for); the dangling-image prune already existed in e2e.yml:422 (the E2E lane keeps it; unchanged, as the PR states); the lock-protocol cross-reference to e2e.yml:419 is accurate; the nightly E2E docker legs are untouched (2-file diff).

Not covered

  • Real dockerized build end-to-end (no image build executed here). Corroborated the two daemon-touching facts directly against this host's real daemon instead: docker info --format '{{.DockerRootDir}}' → /var/lib/docker, and the real helper against it → DISKFLOOR … dir[/var/lib/docker] fs[/] avail_kb[1376919168] floor_kb[8388608] free_inodes[485500488], exit 0. The "8 GiB covers a cold builder stage" sizing is accepted as the author's judgment (env-overridable via DISK_FLOOR_MIN_FREE_KB).
  • Pinned ShellCheck 0.11.0: scripts/lint.js has no linux/aarch64 entry, so the pinned binary cannot run on this host. Corroboration with distro ShellCheck 0.9.0: default check set clean on both arms; with the CI's exact flags (--enable=all …) both arms fail identically on pre-existing style notes (SC2292/SC2250 across the file's [ ] idiom, SC2154 ×2) — an A/A-proven 0.9.0-vs-pinned behavior difference, not attributable to the PR (head adds notes of the same classes only, no new warning class).
  • Full scripts vitest suite — the diff's blast radius is release-workflow.test.js; the targeted suite (2 files, 84 tests) is what was run, on both trees.
  • Step-level logs of run 37374675168 (annotations API 404s with this token scope). Job-level facts corroborated: Integration Tests (Docker) failure, 2026-10-05 21:18→21:46 UTC, schedule run — matching the incident description.
  • The daily sweep's 02:30 UTC schedule and the runner-pool capacity claims (ci: prevent transient ENOSPC on high-concurrency self-hosted runners #10035) are external to this repo.

Methodology

Local maintainer-driven round on Linux aarch64 (Orange Pi), Node 24.14, pnpm per-tree installs via scripts/setup-worktree.js. Head 2560f56274fe27523abadf0a523d8ba6c74c9fbd, base = merge-base f338578520eb765f8605d57dc07af94897345e1d (API baseRefOid was fresh; main has not touched either file since, and git merge-tree --write-tree origin/main <head> merges clean). Base replay copied the head test file onto the base tree verbatim. Mutations applied by exact-string replacement (count-asserted), restored by cp from a sha256-verified pristine backup. The wire harness drove bash --noprofile --norc <real script> in a fixture cwd with stub binaries logging every call, composing the real check-disk-floor.sh; both sides of the wire asserted (stub call log + script output + exit code). Raw logs and all harnesses are in the artifact directory tmp/pr13481-verify-20261007-a/. No prior verification round exists on this PR (comment history checked); this is round 1.


Evidence

01-ab-suite-head-vs-base.png

01-ab-suite-head-vs-base

02-mutation-matrix.png

02-mutation-matrix

03-real-helper-harness.png

03-real-helper-harness

@qwen-code-review-bot qwen-code-review-bot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed. Suggestions are inline.

Not explored to full depth (tool budget reached): "agent reverse-audit (round 6)": shellcheck is not installed in this worktree and I did not fetch it, so the new check_docker_data_root_floor function was never run through the lint lane's ex….

中文说明

已审查。 建议见行内评论。

未探索到全部深度(达到工具调用预算):"agent reverse-audit (round 6)":shellcheck is not installed in this worktree and I did not fetch it, so the new check_docker_data_root_floor function was never run through the lint lane's ex…。

— qwen3.8-max via Qwen Code /review (v0.25.0)

Comment thread .github/scripts/run-release-docker-integration.sh Outdated
Comment thread .github/scripts/run-release-docker-integration.sh Outdated
@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🕐 Review received — an automatic review of the current head is still running, so this round is held until it lands (a push now would cancel it and discard its work, #8888). Your feedback stays queued for the next eligible round.

中文说明

🕐 已收到评审 —— 当前 head 上仍有一轮自动 review 在运行,本轮暂缓(现在推送会取消该 review 并丢弃其工作,#8888)。反馈保持排队,等待下一次可运行的轮次处理。

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🤖 Addressed the latest review feedback (round 3/10). What changed, and what I pushed back on: · 已处理最新评审反馈(第 3/10 轮)。改动内容与我反驳保留之处如下:

Autofix round — PR #13481 (issue #13479)

Commit: b6fae431f8 — fix(release): decouple the docker data-root floor knob and pin the prune guards (#13479)

Both inline findings from the automated reviewer were verified against the code, reproduced, and implemented. No conflicts with main (--conflict false; no merge performed).

Feedback disposition

rc:4202942045 — R1-1: 8 GiB docker build floor shared the workspace gate's env knob — implemented

Verified before changing anything: the pre-fix gate line forwarded the ambient
DISK_FLOOR_MIN_FREE_KB into the docker data-root gate
(DISK_FLOOR_MIN_FREE_KB="${DISK_FLOOR_MIN_FREE_KB:-8388608}"), and this is the
only one of the ten check-disk-floor.sh call sites that sets a floor, so the
documented workspace-floor override silently moved the docker build floor too.
Reproduced with the reviewer's own witness shape: the new harness case passing
DISK_FLOOR_MIN_FREE_KB: '1' in the spawn env was red pre-fix
(floor=1 at both gates).

Change (.github/scripts/run-release-docker-integration.sh): the gate now reads
a dedicated DISK_FLOOR_DOCKER_MIN_FREE_KB knob, still defaulting to 8388608 and
still forwarded to the helper as DISK_FLOOR_MIN_FREE_KB — the helper and its
2 GiB default for the nine other call sites are untouched. A two-line comment
records why the knob must stay separate. This keeps the advertised escape hatch
(the PR's disclosed "8 GiB can in principle fail a build that would have fit"
tradeoff) while decoupling it from the workspace floor.

Change (scripts/tests/release-workflow.test.js):

  • the text pin now asserts DISK_FLOOR_DOCKER_MIN_FREE_KB:-8388608;
  • the harness strips both floor knobs from the ambient ...process.env spread
    (the suite's verdict no longer depends on the runner's environment) and only
    forwards them when a test pins them explicitly;
  • new test keeps the 8 GiB docker floor when the workspace floor knob is set
    passes DISK_FLOOR_MIN_FREE_KB: '1' and asserts floor=8388608 at both gates
    (red pre-fix, green after);
  • new test honors the dedicated docker floor knob pins
    DISK_FLOOR_DOCKER_MIN_FREE_KB=4194304 and asserts it reaches both gates.

rc:4202942062 — R1-2: load-bearing || echo "::warning::…" prune guards unobservable in tests — implemented

Verified: the docker stub's only failure branch was gated on
builder prune + PRUNE_FAILS, so every docker image prune exited 0 in all
tests and the guards on the labelled prune and the new dangling prune had no
witness.

Change (scripts/tests/release-workflow.test.js): the docker stub gained a
separate IMAGE_PRUNE_FAILS branch that fails both docker image prune
invocations (it does not overload PRUNE_FAILS, so the BuildKit cache-prune
test's outcome is unchanged), plus one witness test per guard:
continues with a warning when the dangling image prune fails and
continues with a warning when the labelled image prune fails. Both assert
status 0, the specific ::warning:: text, and that the build still ran.

Notes for the maintainer (PR body only — no code action taken)

  • The PR body advertises the floor override as DISK_FLOOR_MIN_FREE_KB; the
    docker gate's override knob is now DISK_FLOOR_DOCKER_MIN_FREE_KB. A
    one-line body edit would keep the description accurate.
  • @wenshao's verification comment (ic:6029348643, observation ①) already
    flagged a separate body wording nit: the cached-image path does gain one
    pre-vitest gate. Also body-only.

Mutation probes (each performed, then restored; script sha-verified against the pre-round backup)

  1. Reverted the gate line to forwarding DISK_FLOOR_MIN_FREE_KB → red:
    keeps the 8 GiB docker floor…, honors the dedicated docker floor knob,
    and the text-pin test prunes BuildKit cache and gates the docker data root… (3 tests killed). Restored → green.
  2. Removed || echo "::warning::dangling image cleanup failed…" → red:
    continues with a warning when the dangling image prune fails
    (non-zero status, no warning). Restored → green.
  3. Removed || echo "::warning::old CI sandbox image cleanup failed…" → red:
    continues with a warning when the labelled image prune fails
    (non-zero status, no warning). Restored → green.

Verification

  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/release-workflow.test.js — 82 passed (82) on the committed state
  • Reviewer's witness scenario DISK_FLOOR_MIN_FREE_KB=1 npx vitest run … -t 'builds the image only after the cache prune' — passed (was red pre-fix)
  • npm run build — passed (exit 0)
  • npm run typecheck — passed (exit 0)
  • npm run lint — passed (exit 0)
  • npx prettier --check scripts/tests/release-workflow.test.js — passed
  • bash -n .github/scripts/run-release-docker-integration.sh — passed
  • Pre-fix red proof: the two new floor-knob tests failed on the unmodified script with exactly the defect shape (floor=1 under an ambient workspace knob)
  • Not run: the pinned ShellCheck 0.11.0 lane (node scripts/lint.js downloads the binary; outside the trusted command set and not installed on this runner). The change reuses the file's pre-existing ${VAR:-default} env-forward idiom with a renamed variable — no new shell construct. @wenshao's A/A shellcheck corroboration on the pre-round head covers the file's baseline posture.
  • Not run: integration tests (npm run bundle + harness) — the changed behavior (CI shell script + its vitest harness) is not exercised through the bundled CLI.
中文说明

Autofix 本轮处理 — PR #13481(issue #13479)

提交:b6fae431f8 — fix(release): decouple the docker data-root floor knob and pin the prune guards (#13479)

自动化审查的两条行内发现均已对照代码核实、复现并实现。与 main 无冲突(--conflict false,未执行合并)。

反馈处置

rc:4202942045 — R1-1:8 GiB docker 构建下限与工作区门禁共用一个环境变量 —— 已实现

改动前已核实:修复前的门禁行会把环境中的 DISK_FLOOR_MIN_FREE_KB 透传进 docker 数据根目录门禁(DISK_FLOOR_MIN_FREE_KB="${DISK_FLOOR_MIN_FREE_KB:-8388608}"),且这是十处 check-disk-floor.sh 调用点中唯一设置下限的一处,因此有文档记载的工作区下限覆盖开关会静默地同时移动 docker 构建下限。按审查者给出的佐证形态复现:在 spawn 环境传入 DISK_FLOOR_MIN_FREE_KB: '1' 的新脚手架用例在修复前为红(两处门禁均显示 floor=1)。

改动(.github/scripts/run-release-docker-integration.sh):该门禁改为读取专属的 DISK_FLOOR_DOCKER_MIN_FREE_KB 旋钮,默认值仍为 8388608,仍以 DISK_FLOOR_MIN_FREE_KB 透传给 helper——helper 本身及其余九处调用点依赖的 2 GiB 默认值均未改动。一段两行注释记录了该旋钮必须独立的原因。这既保留了宣传中的逃生开关(PR 已披露的“8 GiB 原则上可能误杀本可以完成的构建”这一权衡),又将其与工作区下限解耦。

改动(scripts/tests/release-workflow.test.js):

  • 文本钉断言改为 DISK_FLOOR_DOCKER_MIN_FREE_KB:-8388608;
  • 脚手架会从 ambient 的 ...process.env 展开中剔除两个下限旋钮(套件结论不再依赖运行者的环境变量),只有测试显式钉住时才透传;
  • 新测试 keeps the 8 GiB docker floor when the workspace floor knob is set 传入 DISK_FLOOR_MIN_FREE_KB: '1' 并断言两处门禁仍为 floor=8388608(修复前红、修复后绿);
  • 新测试 honors the dedicated docker floor knob 钉住 DISK_FLOOR_DOCKER_MIN_FREE_KB=4194304 并断言其到达两处门禁。

rc:4202942062 — R1-2:起承重作用的 || echo "::warning::…" 清理保护在测试中不可观察 —— 已实现

已核实:docker stub 唯一的失败分支被限定在 builder prune + PRUNE_FAILS 上,因此所有测试中每一次 docker image prune 都以 0 退出,带标签清理与新增悬空清理上的两处保护都没有见证。

改动(scripts/tests/release-workflow.test.js):docker stub 新增独立的 IMAGE_PRUNE_FAILS 分支,让两次 docker image prune 调用失败(不复用 PRUNE_FAILS,因此 BuildKit 缓存清理测试的结果不变),并为每处保护各加一个见证测试:continues with a warning when the dangling image prune fails 与 continues with a warning when the labelled image prune fails。两者都断言状态为 0、出现对应的 ::warning:: 文本、且构建仍被执行。

给维护者的说明(仅 PR 正文——未做代码改动)

  • PR 正文把下限覆盖开关写作 DISK_FLOOR_MIN_FREE_KB;docker 门禁的覆盖旋钮现为 DISK_FLOOR_DOCKER_MIN_FREE_KB。改正文一行即可保持描述准确。
  • @wenshao 的验证评论(ic:6029348643,观察①)已指出另一处正文措辞问题:缓存镜像路径现在确实多了一次 vitest 前门禁。同样仅涉及正文。

变异探针(逐一执行后恢复;脚本已与本轮开始前备份做 sha 校验)

  1. 将门禁行回退为透传 DISK_FLOOR_MIN_FREE_KB → 变红:keeps the 8 GiB docker floor…、honors the dedicated docker floor knob 以及文本钉测试 prunes BuildKit cache and gates the docker data root…(3 个测试被杀)。恢复 → 绿。
  2. 删除 || echo "::warning::dangling image cleanup failed…" → 变红:continues with a warning when the dangling image prune fails(状态非零且无警告)。恢复 → 绿。
  3. 删除 || echo "::warning::old CI sandbox image cleanup failed…" → 变红:continues with a warning when the labelled image prune fails(状态非零且无警告)。恢复 → 绿。

验证

  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/release-workflow.test.js —— 提交后状态 82 通过(82)
  • 审查者佐证场景 DISK_FLOOR_MIN_FREE_KB=1 npx vitest run … -t 'builds the image only after the cache prune' —— 通过(修复前为红)
  • npm run build —— 通过(exit 0)
  • npm run typecheck —— 通过(exit 0)
  • npm run lint —— 通过(exit 0)
  • npx prettier --check scripts/tests/release-workflow.test.js —— 通过
  • bash -n .github/scripts/run-release-docker-integration.sh —— 通过
  • 修复前红色证明:两个新的下限旋钮测试在未修改的脚本上失败,失败形态与缺陷完全一致(ambient 工作区旋钮下 floor=1)
  • 未运行:固定版本 ShellCheck 0.11.0 通道(node scripts/lint.js 需联网下载二进制;不在受信命令集合内,且本 runner 未安装)。本次改动复用文件中既有的 ${VAR:-default} 环境变量透传写法,仅改变量名,未引入新的 shell 构造。@wenshao 在本轮前 head 上的 A/A shellcheck 旁证已覆盖该文件的基线状态。
  • 未运行:集成测试(npm run bundle + 集成线束)——被改动的行为(CI shell 脚本及其 vitest 脚手架)不经过打包后的 CLI。

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。

🧵 Resolved all 2 selected review thread(s). · 已关闭全部选中的 2 条评审线程。

Re-review when you have a moment. After round 10 this bot stops and leaves the PR for a human. · 有空请复审;第 10 轮后本 bot 停止并将 PR 交给人工。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.25.0

@qwen-code-review-bot qwen-code-review-bot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed. Suggestions are inline.

Not explored to full depth (tool budget reached): "agent 4": running the 11 new spawnSync cases to measure their added suite time rather than reasoning about it from the stub harness..

Deferred under the convergence posture (round 2, not a blocker) — recorded, not requested in this round:

  • .github/scripts/run-release-docker-integration.sh:94 — [review] the three timeout 20m bounds inside the fd-7 host build mutex sum to ~61 min against the E2E lane's 1800 s wait on the same lock, so the guarantee the comment states does not h…
  • .github/scripts/run-release-docker-integration.sh:46 — [review] three distinct gate-skip causes report one cause ('is not a readable directory') and 2>/dev/null discards the daemon diagnostic that would distinguish them
中文说明

已审查。 建议见行内评论。

未探索到全部深度(达到工具调用预算):"agent 4":running the 11 new spawnSync cases to measure their added suite time rather than reasoning about it from the stub harness.。

收敛姿态下延后(第 2 轮,非阻断)——已记录,本轮不要求修改:共 2 条(原文未翻译,列表见上方英文部分)。

— qwen3.8-max via Qwen Code /review (v0.25.0)

Comment thread .github/scripts/run-release-docker-integration.sh Outdated
Comment thread scripts/tests/release-workflow.test.js Outdated
@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🕐 Review received — an automatic review of the current head is still running, so this round is held until it lands (a push now would cancel it and discard its work, #8888). Your feedback stays queued for the next eligible round.

中文说明

🕐 已收到评审 —— 当前 head 上仍有一轮自动 review 在运行,本轮暂缓(现在推送会取消该 review 并丢弃其工作,#8888)。反馈保持排队,等待下一次可运行的轮次处理。

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🤖 Addressed the latest review feedback (round 4/10). What changed, and what I pushed back on: · 已处理最新评审反馈(第 4/10 轮)。改动内容与我反驳保留之处如下:

Round summary — PR #13481 (issue #13479)

Commit: 175d41d6a5 fix(release): decouple the docker inode floor knob and drop duplicated text pins (#13479)

Both round-2 Suggestions were reproduced, implemented, and mutation-verified. No feedback was declined, deferred, or escalated; no base-conflict resolution was needed (--conflict false, no merge performed).

rc:4204967052 — R2-1: inode floor still coupled to the job-start gate knob — RESOLVED

Reproduced before fixing. Added the harness's inode-knob plumbing and two new execution tests first; against the unfixed script they failed exactly as the finding's witness predicts: with DISK_FLOOR_MIN_FREE_INODES=999999999 set ambiently, both docker data-root gate invocations logged inodes=999999999 instead of the unscaled default, and setting a dedicated docker inode knob had no effect (inodes=100000 instead of 200000). All 88 other tests stayed green.

Fix (option B — decouple, not just reword the comment). check_docker_data_root_floor now passes DISK_FLOOR_MIN_FREE_INODES="${DISK_FLOOR_DOCKER_MIN_FREE_INODES:-100000}" alongside the existing KB override, and the comment now names both knobs. The docker default stays at the helper's own default of 100000, so behavior is byte-identical when no knob is set, and the helper check-disk-floor.sh is untouched (its default still serves the other call sites). One deliberate deviation from the finding's option-B sketch: the sketch's inner ${DISK_FLOOR_MIN_FREE_INODES:-...} fallback would have kept the ambient pass-through the finding asks to eliminate and contradicts its own acceptance test, so the KB-consistent form (dedicated knob, else fixed default) was used instead.

Tests. The stub check-disk-floor.sh now also records the inode floor (mirroring the real helper's :-100000 default), and the harness gained ambientDiskFloorInodes / dockerFloorInodes parameters with hermetic env deletes. Two new tests pin both arms: keeps the default docker inode floor when the workspace inode knob is set (ambient knob must not leak into either docker gate) and honors the dedicated docker inode knob. The five pre-existing exact-match gate assertions were extended to the new log format.

Mutation probe. Negating the new branch (script hardcodes DISK_FLOOR_MIN_FREE_INODES="100000", ignoring the docker knob) reddened exactly honors the dedicated docker inode knob (1 failed / 89 passed); the pre-fix pass-through state reddened exactly the two new tests. Each arm of the new branch has its own witness.

rc:4204967056 — R2-2: duplicated text-position pins in the prune/gate text test — RESOLVED

Removed the eight duplicated assertions from prunes BuildKit cache and gates the docker data root before the build and the vitest phase — the gate count, the four gate-ordering assertions, the two indentation assertions, and the DISK_FLOOR_DOCKER_MIN_FREE_KB:-8388608 literal — along with the scriptLines/gateCallLines/cachePruneLine/buildLine/firstVitestLine computations that existed only to feed them. Kept everything through the timeout 60 docker info pin (the finding's lines 2151-2174): the full prune spellings, the prune ordering, and the timeout 20m builder-prune / timeout 60 docker-info / --filter 'until=24h' pins the execution harness cannot see. The surviving coverage is behavioral: gate count and order are pinned by builds the image only after the cache prune and the pre-build disk gate pass and gates the cached-image re-run path before the vitest phase, and the knob spelling by honors the dedicated docker floor knob and keeps the 8 GiB docker floor when the workspace floor knob is set. The deletion is recorded in test-weakening.json with this evidence.

Acceptance probe (M1). Reindenting the pre-build gate call one level (2→4 spaces, a bash no-op) left the whole suite green (90 passed); the probe was then reverted and the suite re-run green.

Note on the review-body deferred items. The two entries under the review's own qwen-review-deferred marker (the ~61-minute bound sum vs the 1800 s lock wait, and the single gate-skip cause string) were deferred by the reviewer itself under the convergence posture and were not requested this round; they are unchanged and remain tracked there.

Verification

All run on this checkout with COREPACK_HOME=/tmp/corepack-home (the runner's default corepack cache path is not writable in this sandbox; the suite's globalSetup honors COREPACK_HOME):

  • npx vitest run --config ./scripts/tests/vitest.config.ts release-workflow.test — 90 passed (2 files), run on the final committed tree.
  • Reproduction run before the fix (test edits only): exactly the 2 new inode tests failed (inodes=999999999 and inodes=100000 received vs expected inodes=100000 / inodes=200000), 88 passed.
  • Mutation probe A (docker inode knob ignored by the script): exactly honors the dedicated docker inode knob failed (89 passed); then restored and re-run green.
  • Mutation probe M1 (pre-build gate reindented 2→4 spaces): 90 passed; then reverted and re-run green.
  • bash -n .github/scripts/run-release-docker-integration.sh — passed.
  • npx prettier --experimental-cli --check scripts/tests/release-workflow.test.js — clean; the commit's pre-commit hook (prettier + eslint on staged files) also passed.
  • npm run build — passed.
  • npm run typecheck — passed.
  • npm run lint — passed.
  • npm run test:scripts (full scripts suite) — 2826 passed, 17 failed in 4 files this round did not touch (verify-capture, web-shell-publish-artifacts, install-script, qwen-triage-workflow; tmux-proxy/packaging/ANSI-rendering cases). Reproduced identically on the unmodified HEAD (same 4 files, same 17 failures: 1+11+1+4), so they are environment-specific to this runner and pre-existing relative to this change; the touched release-workflow suite is fully green in both runs.
中文说明

本轮摘要 — PR #13481(issue #13479)

提交:175d41d6a5 fix(release): decouple the docker inode floor knob and drop duplicated text pins (#13479)

第 2 轮的两条 Suggestion 均已复现、实现并通过变异验证。没有驳回、延后或升级的反馈项;无需解决基线冲突(--conflict false,未做合并)。

rc:4204967052 — R2-1:inode 下限仍与作业启动门禁旋钮耦合 — 已解决

修复前先复现。 先补上测试脚手架的 inode 旋钮管线并新增两个执行测试;在未修复的脚本上,它们精确复现了该发现的见证:环境设置 DISK_FLOOR_MIN_FREE_INODES=999999999 时,两处 docker 数据根目录门禁记录的均为 inodes=999999999 而非未缩放的默认值;设置专属 docker inode 旋钮则无效(得到 inodes=100000 而非 200000)。其余 88 个测试保持绿色。

修复(选项 B——解耦,而非仅改写注释)。 check_docker_data_root_floor 现在在既有 KB 覆盖旁同时传入 DISK_FLOOR_MIN_FREE_INODES="${DISK_FLOOR_DOCKER_MIN_FREE_INODES:-100000}",注释也改为同时点名两个旋钮。docker 侧默认值保持 helper 自身的默认值 100000,因此未设置任何旋钮时行为逐字节不变;helper check-disk-floor.sh 未改动(其默认值仍服务于其余调用点)。与发现中选项 B 草样的一处有意偏差:草样里内层的 ${DISK_FLOOR_MIN_FREE_INODES:-...} 回退会保留该发现要求消除的环境透传,并与其自身的验收测试相矛盾,因此采用了与 KB 一致的形式(专属旋钮,否则固定默认值)。

测试。 stub check-disk-floor.sh 现在同时记录 inode 下限(镜像真实 helper 的 :-100000 默认值),脚手架新增 ambientDiskFloorInodes / dockerFloorInodes 参数并做环境隔离删除。两个新测试分别钉住两个分支:keeps the default docker inode floor when the workspace inode knob is set(环境旋钮不得渗入任一 docker 门禁)与 honors the dedicated docker inode knob。五处既有精确匹配门禁断言扩展到了新的日志格式。

变异探针。 否定新分支(脚本硬编码 DISK_FLOOR_MIN_FREE_INODES="100000"、忽略 docker 旋钮)恰好只让 honors the dedicated docker inode knob 变红(1 失败 / 89 通过);修复前的透传状态恰好只让两个新测试变红。新分支的每一臂都有独立见证。

rc:4204967056 — R2-2:prune/门禁文本测试中的重复文本位置钉断 — 已解决

从 prunes BuildKit cache and gates the docker data root before the build and the vitest phase 中删除了八条重复断言——门禁计数、四条门禁顺序断言、两条缩进断言,以及 DISK_FLOOR_DOCKER_MIN_FREE_KB:-8388608 字面量——连同只为它们服务的 scriptLines/gateCallLines/cachePruneLine/buildLine/firstVitestLine 计算。保留了到 timeout 60 docker info 钉断为止的全部内容(即该发现所说的 2151-2174 行):完整的 prune 拼写、prune 顺序,以及执行脚手架在结构上看不到的 timeout 20m builder-prune / timeout 60 docker-info / --filter 'until=24h' 钉断。存留覆盖为行为式:门禁计数与顺序由 builds the image only after the cache prune and the pre-build disk gate pass 和 gates the cached-image re-run path before the vitest phase 钉住,旋钮拼写由 honors the dedicated docker floor knob 和 keeps the 8 GiB docker floor when the workspace floor knob is set 钉住。该删除已连同上述证据记录在 test-weakening.json 中。

验收探针(M1)。 将构建前门禁调用多缩进一级(2→4 空格,bash 视为无操作)后整个套件保持绿色(90 通过);随后还原探针并复跑为绿。

关于评审正文延后条目的说明。 评审自身 qwen-review-deferred 标记下的两个条目(约 61 分钟的时限总和对比 1800 秒锁等待,以及单一的门禁跳过原因字符串)是评审者在收敛姿态下自行延后的,本轮未要求处理;它们保持原样并继续在那里被跟踪。

验证

全部在本检出上以 COREPACK_HOME=/tmp/corepack-home 运行(本沙箱中 runner 默认的 corepack 缓存路径不可写;套件的 globalSetup 支持 COREPACK_HOME):

  • npx vitest run --config ./scripts/tests/vitest.config.ts release-workflow.test —— 90 通过(2 个文件),在最终提交后的树上运行。
  • 修复前复现运行(仅测试改动):恰好 2 个新 inode 测试失败(实际收到 inodes=999999999 与 inodes=100000,期望 inodes=100000 / inodes=200000),88 通过。
  • 变异探针 A(脚本忽略 docker inode 旋钮):恰好只有 honors the dedicated docker inode knob 失败(89 通过);随后还原并复跑为绿。
  • 变异探针 M1(构建前门禁缩进 2→4 空格):90 通过;随后还原并复跑为绿。
  • bash -n .github/scripts/run-release-docker-integration.sh —— 通过。
  • npx prettier --experimental-cli --check scripts/tests/release-workflow.test.js —— 干净;提交的 pre-commit 钩子(对暂存文件跑 prettier + eslint)同样通过。
  • npm run build —— 通过。
  • npm run typecheck —— 通过。
  • npm run lint —— 通过。
  • npm run test:scripts(完整 scripts 套件)—— 2826 通过,17 失败,分布在 4 个本轮未触碰的文件(verify-capture、web-shell-publish-artifacts、install-script、qwen-triage-workflow;tmux 代理/打包/ANSI 渲染用例)。在未改动的 HEAD 上完整复现(同样 4 个文件、同样 17 个失败:1+11+1+4),因此属于该 runner 的环境特定失败,相对本次改动为既有问题;被触动的 release-workflow 套件在两次运行中均全绿。

🧪 Gate advisory — this round weakened or removed pre-existing tests (machine-measured, not agent-authored):

  • scripts/tests/release-workflow.test.js — net 2 assertion(s) removed

The round recorded evidence for each (below, agent-authored). Weakening is sound only when the pinned behaviour itself was wrong or the coverage demonstrably survives elsewhere — read each reason against the diff. · 本轮弱化或删除了既有测试(门自动测量,非 agent 文本)。下列理由由 agent 撰写:仅当被钉住的行为本身有误、或覆盖确有替代时才成立,请对照 diff 逐条审阅。

  • scripts/tests/release-workflow.test.js: Removed 8 duplicated text-position assertions (gate count, four gate-ordering, two indentation, one floor literal) from the 'prunes BuildKit cache' text test; the pinned behavior survives in named execution tests: gate count/order/position in 'builds the i

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。

🧵 Resolved all 2 selected review thread(s). · 已关闭全部选中的 2 条评审线程。

Re-review when you have a moment. After round 10 this bot stops and leaves the PR for a human. · 有空请复审;第 10 轮后本 bot 停止并将 PR 交给人工。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.25.0

@qwen-code-review-bot qwen-code-review-bot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Partially reviewed — gaps disclosed. Suggestions are inline.

Not reviewed: the executable-script lint — .github/scripts/run-release-docker-integration.sh: shellcheck is not installed.

Not explored to full depth (tool budget reached): "agent reverse-audit (round 2)": could not execute the 15 new spawnSync tests — the review worktree has no node_modules (no node_modules/.bin/vitest ), so their greenness is hand-traced fr…; "agent 2": could not execute scripts/tests/release-workflow.test.js (no node_modules in the review worktree), so the macOS-lane clearance above is read-based rather th…; "agent 6a": could not execute scripts/tests/release-workflow.test.js to confirm the new tests are green — the review worktree has no node_modules , so vitest cannot re….

Deferred under the convergence posture (round 3, not a blocker) — recorded, not requested in this round:

  • .github/scripts/run-release-docker-integration.sh:96 — [review] the builder prune's --filter 'until=24h' excludes the same-day accumulation it exists to reclaim at the 21:00 UTC nightly (sweep is 02:30 UTC), so on the incident host it frees…
  • scripts/tests/release-workflow.test.js:2614 — [review] the dangling-prune and labelled-prune failure tests flip one stub switch, so swapping the two operator-facing ::warning:: texts keeps the suite green and a defect in the labelled prune …
  • scripts/tests/release-workflow.test.js:2520 — [review] 'skips the reclaim and the gate when the image already exists' runs github-hosted, so its not.toContain('disk-floor') is satisfied by the environment guard and cannot fail for the image…

Convergence: round 3 posted 1 inline comment(s), 1 of them reported for the first time; the previous round posted 2 (2 new). Findings keep coming back to the same files: .github/scripts/run-release-docker-integration.sh (findings in round 2; 1 more now). A cluster that keeps producing siblings usually means the fixes are treating instances of a shared root cause — triaging that cause before the next round, or splitting an independent cluster into its own pull request, tends to end the loop faster than fixing them one at a time. (Observation only — nothing was withheld from this review because of this observation.)

中文说明

仅完成部分审查,审查缺口已披露。 建议见行内评论。

未审查(原文为英文):the executable-script lint — .github/scripts/run-release-docker-integration.sh: shellcheck is not installed.

未探索到全部深度(达到工具调用预算):"agent reverse-audit (round 2)":could not execute the 15 new spawnSync tests — the review worktree has no node_modules (no node_modules/.bin/vitest ), so their greenness is hand-traced fr…;"agent 2":could not execute scripts/tests/release-workflow.test.js (no node_modules in the review worktree), so the macOS-lane clearance above is read-based rather th…;"agent 6a":could not execute scripts/tests/release-workflow.test.js to confirm the new tests are green — the review worktree has no node_modules , so vitest cannot re…。

收敛姿态下延后(第 3 轮,非阻断)——已记录,本轮不要求修改:共 3 条(原文未翻译,列表见上方英文部分)。

收敛情况:第 3 轮发布了 1 条行内评论,其中 1 条是首次提出;上一轮发布了 2 条(其中 2 条首次提出)。发现反复回到同一批文件:.github/scripts/run-release-docker-integration.sh(第 2 轮已出过发现,本轮又有 1 条)。一个不断再生兄弟发现的簇,通常意味着逐条修复只在处理同一根因的实例——先定位并处理该根因,或把独立的簇拆成单独的 PR,通常比逐条修复更快结束循环。(仅为观察——本轮评审未因此扣留任何内容。)

— qwen3.8-max via Qwen Code /review (v0.25.0)

fi
# Dedicated knobs, not the job-start gate's DISK_FLOOR_MIN_FREE_KB and
# DISK_FLOOR_MIN_FREE_INODES: one env setting must not move both floors.
DISK_FLOOR_MIN_FREE_KB="${DISK_FLOOR_DOCKER_MIN_FREE_KB:-8388608}" \

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R3-1: The description still names DISK_FLOOR_MIN_FREE_KB as this gate's override, and still says a run that finds the image already built is unaffected. After the last two rounds both statements are false, in the English and the Chinese halves.

Line 51 reads DISK_FLOOR_DOCKER_MIN_FREE_KB, not DISK_FLOOR_MIN_FREE_KB. And the gate at line 119 sits after the fi that closes the image-missing branch, so a cached-image re-run is gated and can fail without building — which this PR's own new tests pin (gates the cached-image re-run path before the vitest phase, blocks the vitest phase when the cached-image host fails the gate). Nothing in the repo documents the new knobs: a tree-wide grep finds them only in this script and its tests, and check-disk-floor.sh:15-16 documents only the two generic ones.

This bites at incident-response time, which is when there is least room to discover it. The Risk section already accepts that the 8 GiB floor "can in principle fail a build that would have fit". When that fires, the release on-call follows the merged description and sets DISK_FLOOR_MIN_FREE_KB=4194304 in the runner/org env, then re-runs. The docker data-root gate is unaffected — line 51 reads the DOCKER knob, still unset, so it stays at 8388608 — and the release stays blocked while the on-call believes the floor was lowered. The value meanwhile flows unmediated into all nine job-start gates that call the helper with no prefix assignment (release.yml:213, ci.yml:627/1149/1528/2101/2323, e2e.yml:297, tui-parity.yml:43/97), raising those workspace floors from the 2 GiB default. And because the description says the cached-image path is unchanged, nobody looks at line 119 — precisely what fails a re-run.

Witness:

real check_disk_floor() body extracted verbatim from script:36-54, driven
against the real check-disk-floor.sh with a stubbed `docker info`:

A0 no knobs                              -> floor_kb[8388608] exit=0
A1 DISK_FLOOR_MIN_FREE_KB=4194304        -> floor_kb[8388608] exit=0
                                            ^ the knob the description names: NO EFFECT
A2 DISK_FLOOR_DOCKER_MIN_FREE_KB=4194304 -> floor_kb[4194304] exit=0
                                            ^ the real knob: flips A1
A4 bare job-start shape (release.yml:213) with DISK_FLOOR_MIN_FREE_KB=4194304
                                         -> floor_kb[4194304] exit=0
                                            ^ the nine other gates DO move

Please update both halves of the description ("How to verify" and "Risk & Scope", EN and ZH) to say that the docker gate reads DISK_FLOOR_DOCKER_MIN_FREE_KB / DISK_FLOOR_DOCKER_MIN_FREE_INODES, that DISK_FLOOR_MIN_FREE_KB / _INODES now govern only the job-start workspace gates, and that a post-build/pre-vitest gate also runs on the cached-image path, so a re-run can fail at the gate without building. The correction must not be made by renaming the helper's own knobs back — check-disk-floor.sh:22-23 (MIN_FREE_KB="${DISK_FLOOR_MIN_FREE_KB:-2097152}" / MIN_FREE_INODES="${DISK_FLOOR_MIN_FREE_INODES:-100000}") serves the nine other job-start call sites at those defaults.

中文说明

PR 描述仍把 DISK_FLOOR_MIN_FREE_KB 写作该门禁的覆盖旋钮,也仍称"并发运行已为同一修订版本构建好镜像时行为完全不变"。经过最近两轮之后,这两句在英文与中文两部分中都不再成立。

第 51 行读取的是 DISK_FLOOR_DOCKER_MIN_FREE_KB,而不是 DISK_FLOOR_MIN_FREE_KB。第 119 行的门禁位于关闭"镜像不存在"分支的 fi 之后,因此发现镜像已缓存的重跑同样会被门禁检查,并可能在完全不构建的情况下失败——本 PR 自己新增的测试已钉住这一点(gates the cached-image re-run path before the vitest phase、blocks the vitest phase when the cached-image host fails the gate)。仓库中没有任何文档记录新旋钮:全树 grep 只在脚本本身及其测试中找到它们,而 check-disk-floor.sh:15-16 只记录了两个通用旋钮。

这个问题会在故障处置时显现,而那正是最没有余裕去发现它的时候。"风险与范围"一节已经接受了 8 GiB 下限"理论上可能误杀本可以完成的构建"这一权衡。当它真的发生时,发布值班同学会照着合并后的描述,在 runner/组织级 env 中设置 DISK_FLOOR_MIN_FREE_KB=4194304 再重跑。docker 数据根目录门禁不受影响——第 51 行读取的是 DOCKER 旋钮,仍未设置,因此保持 8388608——发布依旧被拦住,而值班同学却以为下限已经调低。与此同时,该值会未经中介地流入其余九处以无前缀赋值方式调用 helper 的作业启动门禁(release.yml:213、ci.yml:627/1149/1528/2101/2323、e2e.yml:297、tui-parity.yml:43/97),把它们的工作区下限从 2 GiB 默认值抬高。而且由于描述称镜像已缓存路径不变,没人会去看第 119 行——恰恰是它会让重跑失败。

请更新描述的两个语言版本("如何验证"与"风险与范围"),说明:docker 门禁读取的是 DISK_FLOOR_DOCKER_MIN_FREE_KB / DISK_FLOOR_DOCKER_MIN_FREE_INODES;DISK_FLOOR_MIN_FREE_KB / _INODES 现在只作用于作业启动时的工作区门禁;构建后/vitest 前的门禁在镜像已缓存路径上同样会运行,因此重跑可能在不构建的情况下于门禁处失败。修复时不得把 helper 自身的旋钮改回原名——check-disk-floor.sh:22-23(MIN_FREE_KB="${DISK_FLOOR_MIN_FREE_KB:-2097152}" / MIN_FREE_INODES="${DISK_FLOOR_MIN_FREE_INODES:-100000}")以这些默认值服务于其余九处作业启动调用点。

— qwen3.8-max via Qwen Code /review (v0.25.0)

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified against the current head (175d41d6a5) — the staleness is real, and all three claims check out:

  • run-release-docker-integration.sh:51-53 reads DISK_FLOOR_DOCKER_MIN_FREE_KB / DISK_FLOOR_DOCKER_MIN_FREE_INODES (defaults 8388608 / 100000) and forwards them to the helper as DISK_FLOOR_MIN_FREE_KB / DISK_FLOOR_MIN_FREE_INODES — the knob the description names has no effect on this gate.
  • The post-build gate (run-release-docker-integration.sh:113-120) sits after the fi that closes the image-missing branch, so a cached-image re-run is gated before vitest — pinned by gates the cached-image re-run path before the vitest phase and blocks the vitest phase when the cached-image host fails the gate (scripts/tests/release-workflow.test.js:2535,2566).
  • check-disk-floor.sh:22-23 still reads the generic DISK_FLOOR_MIN_FREE_KB / DISK_FLOOR_MIN_FREE_INODES (defaults 2 GiB / 100000), serving the nine job-start call sites (release.yml:213, ci.yml:627/1149/1528/2101/2323, e2e.yml:297, tui-parity.yml:43/97) — so the helper's own knobs must not be renamed back, exactly as the finding says.

Why this is a reply and not a fix: the requested change is a PR-description edit, and this address-review lane has no PR-body write channel — the workflow sets the body once at PR creation and never edits it afterwards, and the agent side carries no GitHub credentials. A maintainer edit is required. Proposed replacement statements (the body's exact wording may differ slightly from the creation-time quotes; the three semantic corrections are what matter):

  1. Knob sentence (currently "…overridable via DISK_FLOOR_MIN_FREE_KB") → "…overridable via DISK_FLOOR_DOCKER_MIN_FREE_KB (and DISK_FLOOR_DOCKER_MIN_FREE_INODES for the inode floor). The generic DISK_FLOOR_MIN_FREE_KB / DISK_FLOOR_MIN_FREE_INODES knobs now govern only the job-start workspace gates and no longer move the docker data-root floor." · 中文:"……可通过 DISK_FLOOR_DOCKER_MIN_FREE_KB 覆盖(inode 下限用 DISK_FLOOR_DOCKER_MIN_FREE_INODES)。通用的 DISK_FLOOR_MIN_FREE_KB / DISK_FLOOR_MIN_FREE_INODES 现在只作用于作业启动时的工作区门禁,不再影响 docker 数据根目录下限。"
  2. Cached-image sentence (currently "when the image already exists … nothing changes" / "行为完全不变") → "When the image already exists (built by a concurrent run at the same revision) the reclaim and the build are skipped, but the post-build/pre-vitest docker data-root gate still runs — a re-run on a still-saturated host can fail at that gate without building." · 中文:"当镜像已存在(由同一修订版本的并发运行构建)时,回收与构建会跳过,但构建后/vitest 前的 docker 数据根目录门禁仍会运行——在仍然饱和的主机上重跑可能在完全不构建的情况下于该门禁处失败。"
  3. Risk & Scope should carry the same two facts: the gate's overrides are the DISK_FLOOR_DOCKER_* knobs, and the gate also runs on the cached-image path.

Question for the maintainer: could you apply the edit above (happy to iterate on the wording)? The thread stays open until the body is corrected.

中文说明

已对照当前 head(175d41d6a5)核实——描述过时属实,三项断言全部成立:

  • run-release-docker-integration.sh:51-53 读取的是 DISK_FLOOR_DOCKER_MIN_FREE_KB / DISK_FLOOR_DOCKER_MIN_FREE_INODES(默认值 8388608 / 100000),再以 DISK_FLOOR_MIN_FREE_KB / DISK_FLOOR_MIN_FREE_INODES 透传给 helper——描述中点名的旋钮对该门禁无效。
  • 构建后门禁(run-release-docker-integration.sh:113-120)位于关闭"镜像不存在"分支的 fi 之后,因此镜像已缓存的重跑同样会在 vitest 前被门禁检查——已由测试 gates the cached-image re-run path before the vitest phase 与 blocks the vitest phase when the cached-image host fails the gate(scripts/tests/release-workflow.test.js:2535,2566)钉住。
  • check-disk-floor.sh:22-23 仍读取通用的 DISK_FLOOR_MIN_FREE_KB / DISK_FLOOR_MIN_FREE_INODES(默认值 2 GiB / 100000),服务于九处作业启动调用点(release.yml:213、ci.yml:627/1149/1528/2101/2323、e2e.yml:297、tui-parity.yml:43/97)——因此正如该发现所言,helper 自身的旋钮不得改回原名。

为什么这里是回复而不是修复: 请求的改动是 PR 描述编辑,而本 address-review 通道没有 PR 正文写入能力——工作流只在创建 PR 时写入一次正文,之后从不编辑,agent 侧也不持有 GitHub 凭据。需要维护者手动编辑。建议的替换文案如上(英文部分三条),正文实际措辞可能与创建时的引用略有出入,关键是三处语义更正。

请维护者定夺: 能否按上述建议编辑 PR 正文(措辞可继续打磨)?在正文更正前,本线程保持开放。

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R3-1: The description still names DISK_FLOOR_MIN_FREE_KB as this gate's override, in both the English and the Chinese half, but the code reads DISK_FLOOR_DOCKER_MIN_FREE_KB — and a test in this same PR pins that the named knob is deliberately inert. It also still says a run that finds the image already built is unaffected, while the second gate now sits outside the image-missing branch and two tests assert it fires there. This was raised last round and the author's reply confirmed the staleness at 175d41d6a5; the description text is unchanged at this head, so it still stands.

An oncall who follows the description sets DISK_FLOOR_MIN_FREE_KB to lower the docker floor on a saturated host, nothing moves, the job keeps failing at 8 GiB, and they conclude the knob is broken. A reviewer verifying the PR against its own description looks for an unchanged cached-image path, finds a gate firing on it, and either misreads the change as unintended or spends a round asking about it.

Witness:

DISK_FLOOR_MIN_FREE_KB=1        -> OBSERVED floor_kb[8388608]   (the knob the description names)
DISK_FLOOR_DOCKER_MIN_FREE_KB=1 -> OBSERVED floor_kb[1]        (the knob the code reads)

Update the description in both languages: name DISK_FLOOR_DOCKER_MIN_FREE_KB / DISK_FLOOR_DOCKER_MIN_FREE_INODES as this gate's overrides, and replace the "nothing changes when the image is already built" claim with what shipped — the second gate runs outside the image-missing branch, so a cached-image re-run is gated before vitest.

The decoupling the description contradicts is pinned inside this PR: keeps the 8 GiB docker floor when the workspace floor knob is set asserts floor=8388608 while the ambient knob is 1, and gates the cached-image re-run path before the vitest phase asserts the gate fires on the cached path.

中文说明

PR 描述(中英文两部分)仍把 DISK_FLOOR_MIN_FREE_KB 写成这个检查点的覆盖变量,但代码读的是 DISK_FLOOR_DOCKER_MIN_FREE_KB——而且本 PR 自己就有一个测试钉住了“描述里那个变量对本门无效”。描述也仍写着“并发运行已构建好镜像时行为不变”,而第二个门现在位于 image-missing 分支之外,两个测试断言它会在缓存路径上触发。上一轮已提出此问题,作者回复确认在 175d41d6a5 上描述确实过期;本 head 上描述文本未变,因此该发现依然成立。

按照描述去操作的值班同学会设置 DISK_FLOOR_MIN_FREE_KB 来降低饱和主机上的 docker 下限,结果什么都不会变,作业继续以 8 GiB 失败,于是他判断这个开关坏了。对照描述来验证 PR 的评审者会去找“未改动的缓存镜像路径”,却发现该路径上有门在触发,于是要么误读改动意图,要么白花一轮来询问。

证据:

DISK_FLOOR_MIN_FREE_KB=1        -> OBSERVED floor_kb[8388608]   (描述里点名的变量)
DISK_FLOOR_DOCKER_MIN_FREE_KB=1 -> OBSERVED floor_kb[1]        (代码实际读取的变量)

请在两种语言中同步更新描述:把本门的覆盖变量写成 DISK_FLOOR_DOCKER_MIN_FREE_KB / DISK_FLOOR_DOCKER_MIN_FREE_INODES,并把“镜像已存在时行为不变”的说法改成实际落地的行为——第二个门在 image-missing 分支之外运行,因此缓存镜像的重跑在进入 vitest 之前也会被检查。

描述所 contradict 的解耦在本 PR 内已有测试钉住:keeps the 8 GiB docker floor when the workspace floor knob is set 在环境变量为 1 时断言 floor=8388608;gates the cached-image re-run path before the vitest phase 断言缓存路径上门会触发。

— qwen3.8-max via Qwen Code /review (v0.25.0)

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed still stale at this head — and this round widened the gap, so the correction below is current as of this commit. The bot cannot edit the PR body (the workflow writes it only at creation), so this needs a maintainer to paste. Two statements need fixing in both language halves:

  1. The override knobs. The docker data-root gates read DISK_FLOOR_DOCKER_MIN_FREE_KB (pre-build floor, default 8 GiB), DISK_FLOOR_DOCKER_POST_MIN_FREE_KB (pre-vitest floor, default 2 GiB — added this round), and DISK_FLOOR_DOCKER_MIN_FREE_INODES (shared inode floor, default 100000). The job-start gate's DISK_FLOOR_MIN_FREE_KB / DISK_FLOOR_MIN_FREE_INODES move neither docker floor — a test pins exactly that.
  2. The cached-image claim. A re-run that finds the image already built is no longer unaffected: the dangling-image and BuildKit-cache prunes now run on every path, and the pre-vitest gate fires on the cached path before vitest writes container layers.

Suggested replacement (English): "Two disk-floor gates run against the docker data root: a build-sized floor (8 GiB, DISK_FLOOR_DOCKER_MIN_FREE_KB) before the image build and a job-sized floor (2 GiB, DISK_FLOOR_DOCKER_POST_MIN_FREE_KB) before the vitest phase. Both run on every path, including cached-image re-runs, and DISK_FLOOR_DOCKER_MIN_FREE_INODES sets their shared inode floor."

Suggested replacement (中文): “针对 docker 数据根目录运行两道磁盘下限门:构建前按构建体量下限(8 GiB,DISK_FLOOR_DOCKER_MIN_FREE_KB),vitest 阶段前按作业体量下限(2 GiB,DISK_FLOOR_DOCKER_POST_MIN_FREE_KB)。两道门在包括缓存镜像重跑在内的所有路径上都会执行,DISK_FLOOR_DOCKER_MIN_FREE_INODES 设置两者共用的 inode 下限。”

中文说明

已确认在当前 head 上描述仍然过期——而本轮改动进一步拉开了差距,因此下面的更正文本以本次提交为准。机器人无法编辑 PR 描述(工作流只在创建 PR 时写入描述),所以需要维护者手动粘贴。中英文两半中各有两处表述需要修正:

  1. 覆盖变量。 docker 数据根目录的门读取的是 DISK_FLOOR_DOCKER_MIN_FREE_KB(构建前下限,默认 8 GiB)、DISK_FLOOR_DOCKER_POST_MIN_FREE_KB(vitest 前下限,默认 2 GiB——本轮新增)以及 DISK_FLOOR_DOCKER_MIN_FREE_INODES(共用 inode 下限,默认 100000)。作业起始门的 DISK_FLOOR_MIN_FREE_KB / DISK_FLOOR_MIN_FREE_INODES 对两道 docker 门都不起作用——有测试专门钉住这一点。
  2. 缓存镜像的说法。 发现镜像已构建的重跑不再“行为不变”:悬空镜像裁剪与 BuildKit 缓存裁剪现在在所有路径上都会执行,且 vitest 前门会在缓存路径上、vitest 写入容器层之前触发。

建议替换文本(English):“Two disk-floor gates run against the docker data root: a build-sized floor (8 GiB, DISK_FLOOR_DOCKER_MIN_FREE_KB) before the image build and a job-sized floor (2 GiB, DISK_FLOOR_DOCKER_POST_MIN_FREE_KB) before the vitest phase. Both run on every path, including cached-image re-runs, and DISK_FLOOR_DOCKER_MIN_FREE_INODES sets their shared inode floor.”

建议替换文本(中文):“针对 docker 数据根目录运行两道磁盘下限门:构建前按构建体量下限(8 GiB,DISK_FLOOR_DOCKER_MIN_FREE_KB),vitest 阶段前按作业体量下限(2 GiB,DISK_FLOOR_DOCKER_POST_MIN_FREE_KB)。两道门在包括缓存镜像重跑在内的所有路径上都会执行,DISK_FLOOR_DOCKER_MIN_FREE_INODES 设置两者共用的 inode 下限。”

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R3-1: The PR description still names DISK_FLOOR_MIN_FREE_KB as this gate's override, and still says a run that finds the image already built is unaffected. Both statements are false at this head, in the English half and in the Chinese half, and this is the fourth round carrying the entry. The code reads three dedicated knobs (DISK_FLOOR_DOCKER_MIN_FREE_KB at :114, DISK_FLOOR_DOCKER_POST_MIN_FREE_KB at :140, DISK_FLOOR_DOCKER_MIN_FREE_INODES at :59), and the two unlabelled prunes plus the pre-vitest gate now run on every path including cached-image re-runs.

A maintainer deciding whether this change can affect an in-flight or re-run release reads "when a concurrent run already built the image for the same revision, nothing changes" and concludes the cached-image path is untouched, when at this head that path runs two timeout 20m prunes under the shared daemon lock and a 2 GiB data-root gate before vitest, so it can now fail fast where it previously proceeded. An operator who sets DISK_FLOOR_MIN_FREE_KB to relax the 8 GiB build floor sees no effect at all, because the gate reads DISK_FLOOR_DOCKER_MIN_FREE_KB; a test in this same PR pins that the named knob moves neither docker floor.

Witness:

PR description (English half): "a guarded call to check-disk-floor.sh against the Docker data root
                                (8 GiB floor, overridable via DISK_FLOOR_MIN_FREE_KB)"
PR description (English half): "when a concurrent run already built the image for the same revision,
                                nothing changes"
code at HEAD :114   check_docker_data_root_floor "${DISK_FLOOR_DOCKER_MIN_FREE_KB:-8388608}"
code at HEAD :140   check_docker_data_root_floor "${DISK_FLOOR_DOCKER_POST_MIN_FREE_KB:-2097152}"
code at HEAD :59    DISK_FLOOR_MIN_FREE_INODES="${DISK_FLOOR_DOCKER_MIN_FREE_INODES:-100000}" \
code at HEAD :90,:97 both unlabelled prunes now sit OUTSIDE the image-missing branch, which opens at :99
grep -rn 'DISK_FLOOR_DOCKER' --include='*.md' .  ->  no match

The author already drafted the corrected text for both language halves in comment 4211178811 and confirmed the bot cannot edit the PR body, so this needs a maintainer to paste it. Because that is the only route, the entry will keep coming back each round until it is pasted or explicitly declined.

The fix cannot come from the autofix loop this PR runs in: comment 4211178811 records "The bot cannot edit the PR body (the workflow writes it only at creation), so this needs a maintainer to paste." There is no test to add — the correction is to the description text, not to code.

中文说明

[建议] R3-1:PR 描述仍然把 DISK_FLOOR_MIN_FREE_KB 写成这道门的覆盖变量,也仍然写着“发现镜像已构建的运行不受影响”。在当前 head 上这两句都不成立,中英文两半都是,而这已是连续第四轮携带该条目。代码读的是三个专用旋钮(:114 的 DISK_FLOOR_DOCKER_MIN_FREE_KB、:140 的 DISK_FLOOR_DOCKER_POST_MIN_FREE_KB、:59 的 DISK_FLOOR_DOCKER_MIN_FREE_INODES),而且两个不带标签的裁剪与 vitest 前的门现在在包括缓存镜像重跑在内的所有路径上都会执行。

维护者若想判断本改动是否会影响进行中或重跑的发布,读到“当并发运行已为同一修订版本构建好镜像时,行为完全不变”就会以为缓存镜像路径没被动过;而在当前 head 上,该路径会在共享守护进程锁下执行两次 timeout 20m 裁剪,并在 vitest 之前过一道 2 GiB 的数据根目录门,因此它现在可能快速失败,而以前会继续执行。运维若设置 DISK_FLOOR_MIN_FREE_KB 想放宽 8 GiB 构建下限,则完全没有效果,因为门读的是 DISK_FLOOR_DOCKER_MIN_FREE_KB;本 PR 自带的测试正是钉住了“该变量两道 docker 门都不会移动”。

证据:

PR 描述(英文半):"a guarded call to check-disk-floor.sh against the Docker data root
                                (8 GiB floor, overridable via DISK_FLOOR_MIN_FREE_KB)"
PR 描述(英文半):"when a concurrent run already built the image for the same revision,
                                nothing changes"
HEAD 代码 :114   check_docker_data_root_floor "${DISK_FLOOR_DOCKER_MIN_FREE_KB:-8388608}"
HEAD 代码 :140   check_docker_data_root_floor "${DISK_FLOOR_DOCKER_POST_MIN_FREE_KB:-2097152}"
HEAD 代码 :59    DISK_FLOOR_MIN_FREE_INODES="${DISK_FLOOR_DOCKER_MIN_FREE_INODES:-100000}" \
HEAD 代码 :90,:97 两个不带标签的裁剪现已位于 image-missing 分支之外(该分支从 :99 开始)
grep -rn 'DISK_FLOOR_DOCKER' --include='*.md' .  ->  无匹配

作者已在评论 4211178811 中为中英文两半起草了更正文本,并确认机器人无法编辑 PR 描述,因此需要维护者手动粘贴。由于这是唯一途径,在被粘贴或被明确拒绝之前,该条目每轮都会重新出现。

修复无法来自本 PR 所在的 autofix 循环:评论 4211178811 记录了“机器人无法编辑 PR 描述(工作流只在创建时写入描述),因此需要维护者粘贴”。无需新增测试——更正的是描述文本,不是代码。

— qwen3.8-max via Qwen Code /review (v0.25.0)

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Still accurate, and still not actionable from this side: the bot cannot edit the PR body (the workflow writes it only at creation), so the corrected text drafted in comment 4211178811 needs a maintainer to paste it. Leaving this thread open as the standing reminder — please paste the corrected wording (both language halves) when convenient.

中文说明

仍然属实,也仍然无法由本侧处理:机器人无法编辑 PR 描述(工作流只在创建时写入描述),因此需要维护者手动粘贴评论 4211178811 中已起草的更正文本。此线程保持开放,作为持续提醒——方便时请粘贴更正后的文本(中英文两半)。

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🕐 Review received — an automatic review of the current head is still running, so this round is held until it lands (a push now would cancel it and discard its work, #8888). Your feedback stays queued for the next eligible round.

中文说明

🕐 已收到评审 —— 当前 head 上仍有一轮自动 review 在运行,本轮暂缓(现在推送会取消该 review 并丢弃其工作,#8888)。反馈保持排队,等待下一次可运行的轮次处理。

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🤖 Reviewed the latest feedback — no changes needed. Why, point by point: · 已审阅最新反馈——无需改动。逐点说明原因如下:

Autofix review round — no action

No code changes this round. The feedback window contained exactly one actionable finding, and its requested fix is a pull-request-description edit, which this lane cannot perform.

  • Inline comments: 1 — rc:4206990470 (R3-1, Suggestion).
  • Reviews: 1 — rv:5442416748 (COMMENTED, automated reviewer; disclosed coverage gaps only, no requested change).
  • Reviewer-deferred items: 3 entries under the review's own qwen-review-deferred marker (the builder-prune until=24h window, the two prune-warning tests sharing one stub switch, and the github-hosted not.toContain('disk-floor') assertion). The reviewer deferred them under the convergence posture and explicitly did not request them this round; left untouched.
  • Failed checks: none. Base conflict: none (--conflict false; no merge performed).

rc:4206990470 — R3-1: PR description still names the old override knob and claims the cached-image path is unaffected — VERIFIED, escalated to the maintainer

Verified every claim against the current head (175d41d6a5) before deciding:

  • .github/scripts/run-release-docker-integration.sh:51-53 reads DISK_FLOOR_DOCKER_MIN_FREE_KB / DISK_FLOOR_DOCKER_MIN_FREE_INODES — the DISK_FLOOR_MIN_FREE_KB knob the description names has no effect on the docker data-root gate.
  • The post-build/pre-vitest gate (.github/scripts/run-release-docker-integration.sh:113-120) sits after the fi closing the image-missing branch, so a cached-image re-run is gated and can fail without building — pinned by the PR's own tests gates the cached-image re-run path before the vitest phase and blocks the vitest phase when the cached-image host fails the gate (scripts/tests/release-workflow.test.js:2535,2566).
  • .github/scripts/check-disk-floor.sh:22-23 reads the generic knobs at the 2 GiB / 100000 defaults serving the nine job-start call sites (release.yml:213, ci.yml:627/1149/1528/2101/2323, e2e.yml:297, tui-parity.yml:43/97 — all confirmed by grep) — so the finding's negative constraint (do not rename the helper's knobs back) is correct, and no code-side variant of this fix exists.

Disposition: escalated, thread left open, with a ready-to-apply correction. The finding asks to update the "How to verify" and "Risk & Scope" sections of the PR description (EN and ZH). This address-review lane has no PR-body write channel: the workflow sets the body once at PR creation and never edits it afterwards (confirmed against qwen-autofix.yml and autofix-push-and-report.sh, which only push commits and post comments), and the agent side carries no GitHub credentials. Round 3 already surfaced this as a maintainer note; it now arrives as an inline finding. I replied on the thread with the verified evidence and the exact replacement statements for both languages, and asked the maintainer to apply the edit. This was not declined — the finding is real and worth fixing — and not deferred to the post-merge queue, because a stale description matters most before merge.

Verification

  • Code inspection of .github/scripts/run-release-docker-integration.sh (gate function, both call sites, branch structure) — finding's claims confirmed.
  • grep of the nine job-start check-disk-floor.sh call sites across .github/workflows/ — confirmed all invoke the helper with no env prefix.
  • Test-name pins confirmed at scripts/tests/release-workflow.test.js:2535,2566.
  • No build/typecheck/lint/vitest runs needed: the working tree is unchanged this round (no commit produced).
中文说明

Autofix 评审轮次 —— 无需改动

本轮没有代码改动。反馈窗口内只有一条可处理的发现,而其请求的修复是 PR 描述编辑,本通道无法执行。

  • 行内评论: 1 条 —— rc:4206990470(R3-1,Suggestion)。
  • 评审: 1 条 —— rv:5442416748(COMMENTED,自动化评审器;仅披露覆盖缺口,未要求改动)。
  • 评审器自行延后的条目: 评审自身 qwen-review-deferred 标记下的 3 个条目(builder prune 的 until=24h 窗口、两个 prune 告警测试共用一个 stub 开关、github-hosted 上的 not.toContain('disk-floor') 断言)。评审器在收敛姿态下自行延后,并明确表示本轮不要求处理;保持原样。
  • 失败检查: 无。基线冲突: 无(--conflict false,未执行合并)。

rc:4206990470 —— R3-1:PR 描述仍写着旧的覆盖旋钮,并声称镜像已缓存路径不受影响 —— 已核实,升级给维护者

在决定处置前,已对照当前 head(175d41d6a5)核实每一项断言:

  • .github/scripts/run-release-docker-integration.sh:51-53 读取的是 DISK_FLOOR_DOCKER_MIN_FREE_KB / DISK_FLOOR_DOCKER_MIN_FREE_INODES——描述中点名的 DISK_FLOOR_MIN_FREE_KB 旋钮对 docker 数据根目录门禁无效。
  • 构建后/vitest 前门禁(.github/scripts/run-release-docker-integration.sh:113-120)位于关闭"镜像不存在"分支的 fi 之后,因此镜像已缓存的重跑会被门禁检查,且可能在完全不构建的情况下失败——已由本 PR 自己的测试 gates the cached-image re-run path before the vitest phase 与 blocks the vitest phase when the cached-image host fails the gate(scripts/tests/release-workflow.test.js:2535,2566)钉住。
  • .github/scripts/check-disk-floor.sh:22-23 以 2 GiB / 100000 默认值读取通用旋钮,服务于九处作业启动调用点(release.yml:213、ci.yml:627/1149/1528/2101/2323、e2e.yml:297、tui-parity.yml:43/97——均已 grep 确认)——因此该发现的否定约束(不得把 helper 的旋钮改回原名)成立,且不存在任何代码侧的修复变体。

处置:升级,线程保持开放,并附可直接套用的更正文案。 该发现要求更新 PR 描述的"How to verify"与"Risk & Scope"两节(英文与中文)。本 address-review 通道没有 PR 正文写入能力:工作流只在创建 PR 时写入一次正文,之后从不编辑(已对照 qwen-autofix.yml 与 autofix-push-and-report.sh 确认,二者只推送提交与发布评论),agent 侧也不持有 GitHub 凭据。第 3 轮已将此问题作为维护者说明提出;现在它以行内发现的形式到来。我已在线程中回复了核实证据与两种语言的精确替换文案,并请维护者应用该编辑。这不是驳回——发现属实且值得修复——也不是延后到合并后队列,因为过时的描述在合并前影响最大。

验证

  • 对 .github/scripts/run-release-docker-integration.sh 的代码审读(门禁函数、两处调用点、分支结构)——发现的断言全部确认。
  • 对 .github/workflows/ 下九处作业启动 check-disk-floor.sh 调用点的 grep——确认全部以无环境前缀方式调用 helper。
  • 测试名钉住已确认于 scripts/tests/release-workflow.test.js:2535,2566。
  • 无需运行 build/typecheck/lint/vitest:本轮工作树无改动(未产生提交)。

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.25.0

@qwen-code-review-bot qwen-code-review-bot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Partially reviewed — gaps disclosed. Suggestions are inline.

2 Suggestion-level finding(s) this review confirmed are already reported on this PR and are not repeated:

  • aggregate prune-bound budget inside the fd-7 build mutex — already reported (round-2 deferral at .github/scripts/run-release-docker-integration.sh:94; maintainer comment 4191354944)
  • shared prune-failure stub switch in the new tests — already reported (round-3 deferral at scripts/tests/release-workflow.test.js:2614)

Not reviewed: closing-issue enumeration — gh in this environment is older than 2.72.0, so the platform's closing-issue references could not be resolved and a second linked issue cannot be ruled out; issue fidelity was evaluated against #13479, which the PR body names explicitly.

Not reviewed: build-and-test — Test (macos-latest, Node 22.x) and Test (windows-latest, Node 22.x) were skipped in CI; the changed suite ran on Linux only (85 passed), so the new spawn-based cases were never executed on Darwin or Windows.

Not reviewed: the executable-script lint — .github/scripts/run-release-docker-integration.sh: shellcheck is not installed.

Not explored to full depth (tool budget reached): "agent reverse-audit (round 2)": executing the 12 new spawn-based tests on a Darwin lane — no macOS host available here, so that lane was reasoned from bash-3.2 compatibility, the timeout / fl….

Convergence: round 4 posted 6 inline comment(s), 5 of them reported for the first time; the previous round posted 1 (1 new). Findings keep coming back to the same files: .github/scripts/run-release-docker-integration.sh (findings in round 3; 5 more now). The rate of new findings is not falling. A cluster that keeps producing siblings usually means the fixes are treating instances of a shared root cause — triaging that cause before the next round, or splitting an independent cluster into its own pull request, tends to end the loop faster than fixing them one at a time. Batching the remaining fixes and verifying them before the next push, or dropping this PR's reviews to --severity-floor critical, keeps the loop from re-deriving the same set. (Observation only — nothing was withheld from this review because of this observation.)

Mechanism health: this round did not close cleanly, so it withholds the incremental anchor — and the round it recovered had no anchor this round could use either — none at all, one with no certifier, one certified by an identity other than the one this round runs under, or one this round's fetch refused or resolved to the head — so the next review re-reads the whole diff unless recovery grafts an earlier own anchor that the round running it can use onto the complete work list this round leaves behind, and keeps doing so until a round's marker carries an anchor again or a graft lands that the round running it can use. (Stated, not acted on — this changes nothing about what the round posts.)

中文说明

仅完成部分审查,审查缺口已披露。 建议见行内评论。

本轮确认的 2 条建议级发现已在 PR 上报告过,不再重复发布(列表见上方英文部分)。

未审查(原文为英文):closing-issue enumeration — gh in this environment is older than 2.72.0, so the platform's closing-issue references could not be resolved and a second linked issue cannot be ruled out; issue fidelity was evaluated against #13479, which the PR body names explicitly.

未审查(原文为英文):build-and-test — Test (macos-latest, Node 22.x) and Test (windows-latest, Node 22.x) were skipped in CI; the changed suite ran on Linux only (85 passed), so the new spawn-based cases were never executed on Darwin or Windows.

未审查(原文为英文):the executable-script lint — .github/scripts/run-release-docker-integration.sh: shellcheck is not installed.

未探索到全部深度(达到工具调用预算):"agent reverse-audit (round 2)":executing the 12 new spawn-based tests on a Darwin lane — no macOS host available here, so that lane was reasoned from bash-3.2 compatibility, the timeout / fl…。

收敛情况:第 4 轮发布了 6 条行内评论,其中 5 条是首次提出;上一轮发布了 1 条(其中 1 条首次提出)。发现反复回到同一批文件:.github/scripts/run-release-docker-integration.sh(第 3 轮已出过发现,本轮又有 5 条)。新发现的产出速度没有下降。一个不断再生兄弟发现的簇,通常意味着逐条修复只在处理同一根因的实例——先定位并处理该根因,或把独立的簇拆成单独的 PR,通常比逐条修复更快结束循环。把剩余修复攒成一批、验证后再推送,或将本 PR 的评审降到 --severity-floor critical,可以避免循环反复推导同一组发现。(仅为观察——本轮评审未因此扣留任何内容。)

机制健康:本轮未能干净收尾,因而扣留了增量锚点,而它恢复到的那一轮也没有留下本轮可用的锚点——要么完全没有、要么没有认证者、要么由本轮运行身份之外的身份认证、要么被本轮的获取拒绝或解析为头提交——因此下一次评审将重读整个 diff,除非恢复流程把本轮能使用的更早自有锚点嫁接到本轮留下的完整工作清单上;并会一直如此,直到某一轮的标记重新带上锚点,或落地的嫁接能被运行该轮的评审使用。(仅陈述,不据此行动——这不改变本轮发布的任何内容。)

— qwen3.8-max via Qwen Code /review (v0.25.0)

fi
fi
docker image prune --all --force --filter 'label=org.qwen-code.ci.sandbox=true' --filter 'until=24h' || echo "::warning::old CI sandbox image cleanup failed on ${RUNNER_NAME:-this runner}"
timeout 20m docker image prune --all --force --filter 'label=org.qwen-code.ci.sandbox=true' --filter 'until=24h' || echo "::warning::old CI sandbox image cleanup failed on ${RUNNER_NAME:-this runner}"

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R4-1: The reclaim sequence added in this block is the pool's fourth copy of the same docker reclaim policy, and the four copies now disagree on which commands they run, on whether they are time-bounded, and on --keep-storage. The copy that matters is the peer E2E lane's: .github/scripts/run-e2e-tests.sh:73 runs the identical labelled prune with no timeout at all while holding the very same docker-sandbox-build.lock this script takes at line 77. So the bound added here protects one direction of a two-way contention — a wedged daemon GC on the E2E side still starves this lane's 30-minute wait. Nothing asserts the copies agree either: ci-runner-routing.test.mjs:590-592 pins e2e.yml's step text and ecs-runner/qwen-docker-cleanup.test.mjs:73-77 pins the sweep's --keep-storage 30GB, so a future reclaim change has to be hand-applied to three shell files plus one workflow and drift between them passes CI silently.

The contention premise is not speculative. docs/plans/2026-09-03-test-suite-barrel-cost.md:44 records that the hk4 runners still carry the shared ecs-qwen label, and scripts/tests/lint.test.js:229 records an ecs-qwen-labelled job actually executing on ecs-qwen-hk4-19 two days before this incident.

Witness:

run-e2e-tests.sh:73        docker image prune --all --force --filter 'label=...' --filter 'until=24h'   <- NO timeout, inside flock --wait 1800 7
qwen-docker-cleanup.sh:44  docker image prune --force --filter 'until=24h'                             <- NO timeout
qwen-docker-cleanup.sh:50  timeout 20m docker builder prune --all ... --keep-storage 30GB              <- under flock --nonblock
e2e.yml:421-422            both image prunes                                                            <- NO timeout, under flock --nonblock 9
this diff :83,:87,:96      timeout 20m x3, NO --keep-storage                                            <- the fourth copy

The in-scope minimum is to bound the peer lane's twin — apply the same timeout to run-e2e-tests.sh:73 so both holders of the mutex are bounded. The fuller fix is one shared reclaim helper taking the label filter and an optional reserve, called from all three scripts and the workflow; that is a cross-lane CI-infra refactor and is reasonably a follow-up issue rather than a change here. If neither lands, a comment recording why the release lane is the only one that needs these guards would stop the next reader treating the omission as an oversight.

A shared helper must keep --keep-storage per call site rather than folding the flag sets into one: ecs-runner/qwen-docker-cleanup.sh:50-51 deliberately retains a 30 GB reserve while this line deliberately omits it, because a reserve reclaims nothing on a host holding less cache than the reserve.

If the minimal variant lands, an assertion in scripts/tests/e2e-workflow.test.js — which already reads the E2E script as e2eRunScript at line 13 — that the labelled prune in run-e2e-tests.sh carries a timeout bound is the test to add; nothing pins that today, so it is red before the fix and green after. Please confirm by removing the bound and watching it fail.

中文说明

本块新增的回收序列是这个 runner 池上同一套 docker 回收策略的第四份副本,而四份副本在“执行哪些命令”“是否有时间上限”“是否带 --keep-storage”三个维度上都不一致。最关键的一份是对端 E2E 通道:.github/scripts/run-e2e-tests.sh:73 在持有本脚本第 77 行同一把 docker-sandbox-build.lock 的情况下,运行完全相同的带标签镜像裁剪,却完全没有 timeout。因此这里加的上限只保护了双向争用中的一个方向——E2E 侧守护进程 GC 卡住时,仍然会饿死本通道 30 分钟的等待。也没有任何测试断言这些副本彼此一致:ci-runner-routing.test.mjs:590-592 钉的是 e2e.yml 的步骤文本,ecs-runner/qwen-docker-cleanup.test.mjs:73-77 钉的是日常清理的 --keep-storage 30GB,所以将来任何回收策略变更都要手工同步到三个 shell 文件加一个 workflow,而它们之间的漂移会静默通过 CI。

争用前提并非猜测:docs/plans/2026-09-03-test-suite-barrel-cost.md:44 记录 hk4 的 runner 仍带共享的 ecs-qwen 标签;scripts/tests/lint.test.js:229 记录了一个带 ecs-qwen 标签的作业确实在 ecs-qwen-hk4-19 上执行过,时间就在本次事故前两天。

证据:

run-e2e-tests.sh:73        docker image prune --all --force --filter 'label=...' --filter 'until=24h'   <- 无 timeout,处于 flock --wait 1800 7 之内
qwen-docker-cleanup.sh:44  docker image prune --force --filter 'until=24h'                             <- 无 timeout
qwen-docker-cleanup.sh:50  timeout 20m docker builder prune --all ... --keep-storage 30GB              <- 在 flock --nonblock 之下
e2e.yml:421-422            两个镜像裁剪                                                                 <- 无 timeout,在 flock --nonblock 9 之下
本 diff :83,:87,:96        timeout 20m x3,无 --keep-storage                                            <- 第四份副本

最小改动范围的做法是给对端那份也加上限:对 run-e2e-tests.sh:73 施加同样的 timeout,让互斥锁的两个持有者都受限。更完整的做法是抽出一个共享回收辅助脚本,接受标签过滤器和可选的保留量参数,由三个脚本和 workflow 共同调用;那属于跨通道的 CI 基础设施重构,放到后续 issue 比放在本 PR 更合适。如果两者都不做,建议加一行注释说明为什么只有发布通道需要这些防护,避免后来的读者把这种缺失当成疏漏。

共享辅助脚本必须把 --keep-storage 保留为逐调用点参数,而不是把两组 flag 合并:ecs-runner/qwen-docker-cleanup.sh:50-51 有意保留 30 GB 储备,而本行有意省略它,因为在缓存量小于储备量的主机上,储备参数什么都回收不了。

如果采用最小方案,需要补的测试在 scripts/tests/e2e-workflow.test.js(它在第 13 行已把 E2E 脚本读作 e2eRunScript):断言 run-e2e-tests.sh 中带标签的裁剪携带 timeout 上限。目前没有任何测试钉住这一点,所以修复前它是红的、修复后是绿的。请通过删掉该上限、确认测试变红来验证。

— qwen3.8-max via Qwen Code /review (v0.25.0)

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified before deciding: run-e2e-tests.sh does run the identical labelled prune with no timeout while holding the same docker-sandbox-build.lock this script takes — the one-direction-bound gap is real.

Deferred to the follow-up queue rather than implemented here: bounding the E2E lane means editing the shared CI script every PR's E2E shards run on, which is CI machinery outside this release-lane PR's footprint — the same call your fuller-fix variant makes. The finding, the minimal fix, and the suggested e2e-workflow.test.js pin are recorded in the deferred-findings issue so they survive the merge.

What did land here, as your fallback option: this script now records the asymmetry next to the mutex — every daemon call it makes is time-bounded so a wedged GC cannot starve a peer lane's 30-minute wait, and the E2E lane's unbounded copy of the same prune under the same lock is noted as a cross-lane change tracked separately.

中文说明

决定前已验证:run-e2e-tests.sh 确实在持有本脚本同一把 docker-sandbox-build.lock 的情况下,运行完全相同的带标签镜像裁剪且不带 timeout——“只约束了一个方向”的缺口属实。

本次不直接实现,而是转入后续跟进队列:给 E2E 通道加上限意味着改动每个 PR 的 E2E 分片都在运行的共享 CI 脚本,属于本发布通道 PR 范围之外的 CI 设施——与您给出的完整修复方案(共享回收辅助脚本)同属跨通道改动。该发现、最小修复方案以及建议补在 e2e-workflow.test.js 的钉住测试都已记录在延期发现 issue 中,合并后不会丢失。

按您给出的兜底方案,本脚本已在互斥锁旁记录了这种不对称:本脚本的每次守护进程调用都带时间上限,卡死的 GC 不会饿死对端通道 30 分钟的锁等待;E2E 通道在同一锁下的同款裁剪仍无上限,注释中已注明这是另行跟踪的跨通道改动。

Comment thread .github/scripts/run-release-docker-integration.sh Outdated
exec 8>&-
fi

# Run 37374675168 actually died in the vitest phase, ~15 minutes after the

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R4-3: Two point samples now cover three heavy phases, and the phase the incident actually died in is the one nothing measures. This gate is taken once, before two back-to-back vitest docker phases, and cleanup_release_containers is on the EXIT trap, so phase-1 containers and their writable layers are still on the data root when phase 2 starts. The companion instrumentation the repo attaches to every other disk-floor gate is absent from this job: ci.yml runs a 10-second df sampler loop next to its gate and uploads the sample file on failure, and integration_docker in release.yml has neither.

A host that clears this sample with just over 8 GiB free can cross zero during the cli phase — container writable layers, transcripts and the runner's own diagnostic log all land on that filesystem — and the worker dies on ENOSPC mid-phase with no step conclusion, so even the always() container-cleanup step never runs and the auto-filed issue again carries no actionable signal. That is the identical shape as run 37374675168. This is not covered by the PR's out-of-scope statement: in-lane instrumentation is neither an end-to-end lane validation nor a host-level capacity control, and #10394 already established it as the companion to every disk-floor gate in ci.yml.

Witness:

DFSAMPLE occurrences:  .github/workflows/ci.yml = 6 (L636, L765, L780, L1158, L1708, L1890)
                       .github/workflows/release.yml = 0
sed -n '530,585p' release.yml | grep -E 'DFSAMPLE|df -|upload-artifact|if: .*failure'  ->  NONE FOUND
only upload-artifact in release.yml is at :340, owned by quality_build
  (job headers: integration_docker = 530, audio_capture_prebuilds = 583)
ci.yml:632-643 is the pattern: sample_disk + ( while sleep 10; do sample_disk; done ) & + trap
  dual sink at :639-640
script:21  trap cleanup_release_containers EXIT     gate at :119, vitest at :122 and :123

Attach the sampler #10394 established to this lane: start a 10-second df loop over "$docker_root" and "${RUNNER_TEMP}" before the first npx vitest, stop it from the existing trap, and add an if: failure() actions/upload-artifact step for the sample file to integration_docker. A third gate sample between the two vitest invocations would also help, but the sampler is what turns a recurrence into a diagnosable one.

Each sample must go to the job log as well as to the file — ci.yml:780 does exactly echo "$sample"; echo "$sample" >> "$DISK_SAMPLES" 2>/dev/null || true — because in the incident the worker itself crashed, the step got no conclusion and even the always() step never ran, so a file plus an if: failure() upload alone would have been lost the same way.

A new case in scripts/tests/release-workflow.test.js in the style of the existing pin at line 349 (gates every pool-routed job on a disk floor before its heavy steps) is the test to add: assert releaseYaml.jobs.integration_docker.steps contains a step whose run starts the sampler loop and an if: failure() upload-artifact step whose path is the sample file. Deleting either must redden it — please confirm by removing one and re-running.

中文说明

现在用两个时间点采样去覆盖三个重负载阶段,而事故真正丧生的那个阶段恰恰没有任何测量。这个门只取一次样,位于两个连续执行的 vitest docker 阶段之前;而 cleanup_release_containers 挂在 EXIT trap 上,所以第一阶段产生的容器及其可写层在第二阶段开始时仍留在数据根目录上。仓库为其它每一个磁盘下限门配备的伴生观测手段在本作业中完全缺席:ci.yml 在其门旁边运行每 10 秒一次的 df 采样循环,并在失败时上传采样文件,而 release.yml 的 integration_docker 两者都没有。

一台以略高于 8 GiB 的余量通过本次采样的主机,可能在 cli 阶段期间归零——容器可写层、测试记录以及 runner 自身的诊断日志都写在同一个文件系统上——于是 worker 在阶段中途因 ENOSPC 死亡,步骤没有任何结论,连 always() 的容器清理步骤都不会运行,自动提交的 issue 再次不带任何可操作信号。这与运行 37374675168 的形态完全相同。这不在本 PR 的“超出范围”声明覆盖之内:通道内观测既不是端到端通道验证,也不是主机级容量控制,而且 #10394 已经把它确立为 ci.yml 中每个磁盘下限门的伴生措施。

证据:

DFSAMPLE 出现次数:  .github/workflows/ci.yml = 6(L636、L765、L780、L1158、L1708、L1890)
                     .github/workflows/release.yml = 0
sed -n '530,585p' release.yml | grep -E 'DFSAMPLE|df -|upload-artifact|if: .*failure'  ->  NONE FOUND
release.yml 中唯一的 upload-artifact 在 :340,属于 quality_build
  (作业起始行:integration_docker = 530,audio_capture_prebuilds = 583)
ci.yml:632-643 即该模式:sample_disk + ( while sleep 10; do sample_disk; done ) & + trap
  双写形式在 :639-640
script:21  trap cleanup_release_containers EXIT     门在 :119,vitest 在 :122 与 :123

请把 #10394 确立的采样器接到本通道上:在第一个 npx vitest 之前启动一个每 10 秒对 "$docker_root" 和 "${RUNNER_TEMP}" 取样的 df 循环,由现有 trap 停止它,并为 integration_docker 增加一个 if: failure() 的 actions/upload-artifact 步骤上传采样文件。在两个 vitest 调用之间再加一次门采样也有帮助,但真正把“再次发生”变成“可诊断”的是采样器。

每个采样必须同时写入作业日志和文件——ci.yml:780 正是 echo "$sample"; echo "$sample" >> "$DISK_SAMPLES" 2>/dev/null || true——因为事故中 worker 自身崩溃、步骤没有结论、连 always() 步骤都未运行,所以只有文件加 if: failure() 上传同样会丢失。

需要补的测试在 scripts/tests/release-workflow.test.js,可参照第 349 行已有的钉法(gates every pool-routed job on a disk floor before its heavy steps):断言 releaseYaml.jobs.integration_docker.steps 中包含一个 run 启动采样循环的步骤,以及一个 if: failure() 的 upload-artifact 步骤且其 path 为采样文件。删除其中任一项都必须让测试变红——请通过删掉一项并重跑来确认。

— qwen3.8-max via Qwen Code /review (v0.25.0)

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified before deciding: release.yml's integration_docker job has no df sampler and no failure upload — ci.yml carries six DFSAMPLE loops, this workflow has none, and its only upload-artifact step belongs to quality_build. The gap is real and matches the incident shape (worker died mid-phase, no step conclusion, even the always() cleanup never ran).

Deferred to the follow-up queue rather than implemented here: attaching the #10394 sampler means adding steps to a workflow this PR never touches — CI machinery outside this PR's footprint. The full design you specified (10-second loop over the docker root and RUNNER_TEMP, dual job-log/file sink per ci.yml:780, if: failure() upload-artifact step, plus the new pin in the style of the line-349 test) is recorded in the deferred-findings issue so it survives the merge.

The secondary ask — a third gate sample between the two vitest invocations — was considered and not added: a point sample there does not diagnose the incident shape (only the streamed sampler survives a worker crash), and each added gate is flake surface on the lane. It rides along in the same deferral record.

中文说明

决定前已验证:release.yml 的 integration_docker 作业确实没有 df 采样器、也没有失败上传——ci.yml 有六处 DFSAMPLE 循环,本工作流一处没有,唯一的 upload-artifact 步骤属于 quality_build。缺口属实,且与事故形态吻合(worker 在阶段中途死亡,步骤没有结论,连 always() 清理步骤都未运行)。

本次不直接实现,而是转入后续跟进队列:接入 #10394 确立的采样器需要向本 PR 从未改动过的工作流添加步骤,属于本 PR 范围之外的 CI 设施。您指定的完整设计(对 docker 根目录与 RUNNER_TEMP 的 10 秒循环、按 ci.yml:780 同时写作业日志与文件的双写、if: failure() 的 upload-artifact 步骤,以及仿照第 349 行钉法的新测试)已记录在延期发现 issue 中,合并后不会丢失。

次要建议——在两个 vitest 调用之间再加一次门采样——经考虑后未采纳:在那里做一次点采样无法诊断事故形态(worker 崩溃时只有流式采样能留存),而且每增加一道门都会给通道带来抖动面。该点已一并记入同一条延期记录。

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R4-3: Still stands. Two point samples now cover three heavy phases, and the phase the incident actually died in is the one nothing measures. This gate is taken once, before two back-to-back vitest docker phases, and cleanup_release_containers is on the EXIT trap, so phase-1 containers and their writable layers are still on the data root when phase 2 starts. The companion instrumentation the repo attaches to every other disk-floor gate is still absent from this job: ci.yml runs a 10-second df sampler loop next to its gate and uploads the sample file on failure, and integration_docker in release.yml has neither.

A host that clears this sample with just over 2 GiB free can cross zero during the cli phase — container writable layers, transcripts and the runner's own diagnostic log all land on that filesystem — and the worker dies on ENOSPC mid-phase with no step conclusion, so even the always() container-cleanup step never runs and the auto-filed issue again carries no actionable signal. That is the identical shape as run 37374675168, which is what this PR exists to make diagnosable. A point sample cannot survive a worker crash; only a streamed sampler can.

Witness (measured in the review worktree at head 46c54cfbe2):

grep -c DFSAMPLE .github/workflows/ci.yml        -> 6
grep -c DFSAMPLE .github/workflows/release.yml   -> 0
grep -c 'while sleep 10' .github/workflows/release.yml -> 0
upload-artifact steps inside integration_docker  -> 0
   (release.yml's single upload-artifact belongs to quality_build)

You verified this gap in comment 4211178534 and deferred it to the follow-up queue because the remedy adds steps to a workflow this PR never touches, which is a reasonable scoping call and is recorded in the deferred-findings issue. Nothing landed this round, so the entry carries forward rather than closing — the deferral record is now the only thing carrying it. The design you specified there is the fix: a 10-second df loop over "$docker_root" and "${RUNNER_TEMP}" started before the first npx vitest, stopped from the existing trap, plus an if: failure() actions/upload-artifact step for the sample file on integration_docker.

Each sample must go to the job log as well as to the file, which is what ci.yml:780 does — echo "$sample"; echo "$sample" >> "$DISK_SAMPLES" 2>/dev/null || true — because in the incident the worker itself crashed, the step got no conclusion and even the always() step never ran, so a file plus an if: failure() upload alone would have been lost the same way. A new case in the style of the existing pin at release-workflow.test.js:349 (gates every pool-routed job on a disk floor before its heavy steps) is the test to add: assert releaseYaml.jobs.integration_docker.steps contains a step whose run starts the sampler loop and an if: failure() upload-artifact step whose path is the sample file, then delete each in turn and confirm it reds.

中文说明

[建议] R4-3:仍然存在。现在是两次点采样覆盖三个重载阶段,而事故真正死掉的那个阶段没有任何东西在测量。这道门只取一次样,位于两个连续的 vitest docker 阶段之前,而 cleanup_release_containers 挂在 EXIT trap 上,因此阶段一的容器及其可写层在阶段二开始时仍然在数据根目录上。仓库为其他每一道磁盘下限门配备的伴生指标采集,在本作业仍然缺席:ci.yml 在其门旁跑一个 10 秒的 df 采样循环并在失败时上传采样文件,而 release.yml 的 integration_docker 两者都没有。

一台以刚好超过 2 GiB 的剩余空间通过本次采样的主机,可能在 cli 阶段降到零——容器可写层、日志输出以及 runner 自身的诊断日志都写在那个文件系统上——于是 worker 在阶段中途因 ENOSPC 死掉,步骤没有任何结论,连 always() 的容器清理步骤都不会运行,自动提交的 issue 再次不带任何可操作信号。这与运行 37374675168 的形态完全一致,而让它可诊断正是本 PR 的目的。点采样无法在 worker 崩溃后存活,只有流式采样可以。

证据(在 head 46c54cfbe2 的评审 worktree 中实测):

grep -c DFSAMPLE .github/workflows/ci.yml        -> 6
grep -c DFSAMPLE .github/workflows/release.yml   -> 0
grep -c 'while sleep 10' .github/workflows/release.yml -> 0
integration_docker 内的 upload-artifact 步骤  -> 0
   (release.yml 唯一的 upload-artifact 属于 quality_build)

你已在评论 4211178534 中验证了该缺口并转入后续跟进队列,理由是补救措施需要向本 PR 从未改动的工作流添加步骤——这是合理的范围判断,也已记入延期发现 issue。本轮没有任何改动落地,因此该条目向前携带而不是关闭——现在延期记录是唯一携带它的东西。你在那里指定的设计就是修复方案:在第一次 npx vitest 之前启动一个针对 "$docker_root" 与 "${RUNNER_TEMP}" 的 10 秒 df 循环,由现有 trap 停止,并为 integration_docker 添加一个 if: failure() 的 actions/upload-artifact 步骤上传采样文件。

每次采样必须同时写入作业日志与文件,也就是 ci.yml:780 的做法——echo "$sample"; echo "$sample" >> "$DISK_SAMPLES" 2>/dev/null || true——因为在那次事故中 worker 本身崩溃、步骤没有结论、连 always() 步骤都未运行,所以仅靠文件加 if: failure() 上传会以同样的方式丢失。要补的测试可仿照 release-workflow.test.js:349 现有的钉法(gates every pool-routed job on a disk floor before its heavy steps):断言 releaseYaml.jobs.integration_docker.steps 包含一个 run 启动采样循环的步骤,以及一个 if: failure() 且 path 为采样文件的 upload-artifact 步骤,然后逐个删除并确认它变红。

— qwen3.8-max via Qwen Code /review (v0.25.0)

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified and still deferred to the follow-up queue: the remedy (a 10-second df sampler over the docker data root and RUNNER_TEMP started before the first npx vitest, stopped from the existing EXIT trap, each sample echoed to the job log as well as the file — the ci.yml:780 pattern, because the incident's worker crash loses a file-only artifact — plus an if: failure() upload-artifact step) adds steps to .github/workflows/release.yml, which this PR never touches. Expanding into a new workflow file mid-review is exactly the scope drift the footprint boundary exists to prevent. The deferral is recorded in the deferred-findings issue so it survives the merge; the design above is the fix to schedule from there.

中文说明

已核实,仍延后至跟进队列:补救方案(在第一次 npx vitest 之前启动一个针对 docker 数据根目录与 RUNNER_TEMP 的 10 秒 df 采样循环,由现有 EXIT trap 停止,每条采样同时写入作业日志与文件——即 ci.yml:780 的做法,因为事故中 worker 崩溃会丢失仅存文件的产物——外加一个 if: failure() 的 upload-artifact 步骤)需要向 .github/workflows/release.yml 添加步骤,而本 PR 从未改动该文件。在评审中途扩展到新的工作流文件,正是足迹边界要防止的范围漂移。该延后已记录在延期发现 issue 中,合并后依然保留;上述方案就是应从那里排期的修复。

Comment thread .github/scripts/run-release-docker-integration.sh
Comment thread .github/scripts/run-release-docker-integration.sh Outdated
@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🕐 Review received — an automatic review of the current head is still running, so this round is held until it lands (a push now would cancel it and discard its work, #8888). Your feedback stays queued for the next eligible round.

中文说明

🕐 已收到评审 —— 当前 head 上仍有一轮自动 review 在运行,本轮暂缓(现在推送会取消该 review 并丢弃其工作,#8888)。反馈保持排队,等待下一次可运行的轮次处理。

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🤖 Addressed the latest review feedback (round 5/10). What changed, and what I pushed back on: · 已处理最新评审反馈(第 5/10 轮)。改动内容与我反驳保留之处如下:

Round summary

All six findings from the round-4 review were triaged source-blind: three were reproduced with the suite's own stub harness and fixed in code (R4-2, R4-4, R4-5), two were verified as real and deferred to the follow-up queue because their fixes live on CI machinery outside this PR's footprint (R4-1, R4-3), and one needs a maintainer to edit the PR body (R3-1). Commit: fix(release): reclaim on the cached path and size the pre-vitest floor separately (#13479).

Fixed in code

  • R4-4 (cached-image re-run reached the pre-vitest gate with no reclaim) — reproduced on the pre-fix script: the cached path logged zero prune lines before the gate, and the two not.toContain assertions in the cached-path tests pinned that gap as intended behavior. Fix: the dangling-image and BuildKit-cache prunes moved out of the image-missing branch so they run on every path before any gate samples the data root; the labelled --all prune stays inside the build mutex (the pre-existing pin expect(imagePrune).toBeGreaterThan(buildLock) still holds). The moved prunes run under the shared daemon lock but outside the build mutex — the daily host sweep runs both with no lock at all — and keep their timeout 20m bounds.
  • R4-5 (the pre-vitest gate re-charged the build budget) — check_docker_data_root_floor now takes the free-space floor as an argument: the pre-build call keeps the build-sized 8 GiB (DISK_FLOOR_DOCKER_MIN_FREE_KB); the pre-vitest call charges the repo's calibrated job floor of 2 GiB (check-disk-floor.sh's own default) via the new DISK_FLOOR_DOCKER_POST_MIN_FREE_KB, so an operator can tune the two checkpoints independently; the inode floor stays shared. A host admitted by the first gate with less than 8 GiB plus the build's own consumption now builds, passes the second gate, and runs both vitest phases instead of dying after the build it already paid for.
  • R4-2 (the dangling-prune comment's justification was measurably false) — rewritten to the verified facts: this lane passes --no-prune to the build, and until on images is creation age, so the line takes dangling images older than a day and a same-day dangler waits for the daily 02:30 UTC sweep. Comment-only; no test pins comment text and the suite held its green baseline.

Deferred to the follow-up queue (verified real, fix outside this PR's footprint)

  • R4-1 (the E2E lane runs the same labelled prune unbounded under the same build mutex) — verified against run-e2e-tests.sh. Bounding it means editing the shared CI script every PR's E2E shards run on; recorded in the deferred-findings issue together with the suggested e2e-workflow.test.js pin. As the finding's own fallback, this script now carries a comment recording the asymmetry so the omission is not mistaken for an oversight. Thread replied, left open.
  • R4-3 (no disk sampler on release.yml's integration_docker) — verified: ci.yml carries six DFSAMPLE loops, release.yml none. Attaching the ci: gate heavy jobs on a disk floor and persist pressure samples #10394 sampler adds steps to a workflow this PR never touches; recorded in the deferred-findings issue with the full design. The secondary mid-phase gate sample was not added: a point sample does not survive the incident's worker crash the way the streamed sampler does, and each added gate is flake surface. Thread replied, left open.

Needs a maintainer

  • R3-1 (PR description still names DISK_FLOOR_MIN_FREE_KB and claims cached-image runs are unaffected) — still stale at this head, and this round widened the gap (new post-floor knob). The bot cannot edit the PR body — the workflow writes it only at creation — so the thread reply carries the exact replacement sentences in both languages for a maintainer to paste. Thread replied, left open.

Effect on earlier deferrals

The round-2 deferral (aggregate prune-bound budget inside the fd-7 build mutex) shrinks: only the labelled prune and the gate's bounded docker info remain inside the mutex. The round-3 deferral (shared prune-failure stub switch in the new tests) is untouched.

Verification

  • bash -n .github/scripts/run-release-docker-integration.sh — passed
  • npx vitest run scripts/tests/release-workflow.test.js — 85 passed at baseline; 87 passed after the fix (two new tests: the post-floor knob, and the host-admitted-between-the-floors case)
  • Mutation probes (each applied, run against the focused suite, then restored and re-run to green):
    • dangling prune gated on image-absence → the 2 cached-path tests failed (R4-4 witness)
    • builder prune gated on image-absence → the same 2 tests failed (R4-4 witness)
    • pre-vitest floor restored to the shared 8 GiB default → 7 floor tests failed (R4-5 witness)
    • pre-vitest floor hardcoded (env knob dropped) → the new post-knob test failed (knob witness)
  • npm run build — passed
  • npm run typecheck — passed
  • npm run lint — passed (the pre-commit hook's prettier/eslint pass also ran clean on both files)
  • npx prettier --check .github/scripts/run-release-docker-integration.sh scripts/tests/release-workflow.test.js — passed
  • Integration tests not run: the change is release-lane shell scripting exercised by the stub-harness suite; the bundled CLI never touches it
  • npm run generate:settings-schema not run: no settings source changed
中文说明

本轮总结

第 4 轮评审的六条发现全部按“来源不加分”的原则分诊:三条经本套件自带的 stub 测试装置复现后在代码中修复(R4-2、R4-4、R4-5);两条验证属实,但因修复落在本 PR 范围之外的 CI 设施上而转入后续跟进队列(R4-1、R4-3);一条需要维护者编辑 PR 描述(R3-1)。提交:fix(release): reclaim on the cached path and size the pre-vitest floor separately (#13479)。

代码中已修复

  • R4-4(缓存镜像重跑在没有任何回收的情况下到达 vitest 前门)——已在修复前的脚本上复现:缓存路径在门之前没有任何裁剪调用,而缓存路径测试中的两条 not.toContain 断言把这个缺口钉成了预期行为。修复:悬空镜像裁剪与 BuildKit 缓存裁剪移出 image-missing 分支,在所有路径上、任何门采样数据根目录之前执行;带标签的 --all 裁剪留在构建互斥锁内(既有钉法 expect(imagePrune).toBeGreaterThan(buildLock) 仍然成立)。移出的两个裁剪在共享守护进程锁之下、构建互斥锁之外运行——每日主机清理对这两个裁剪完全不加锁,已有先例——并保留各自的 timeout 20m 上限。
  • R4-5(vitest 前门重复收取构建预算)——check_docker_data_root_floor 改为以参数接收空闲空间下限:构建前调用保持构建体量的 8 GiB(DISK_FLOOR_DOCKER_MIN_FREE_KB);vitest 前调用通过新增的 DISK_FLOOR_DOCKER_POST_MIN_FREE_KB 收取仓库标定的作业下限 2 GiB(check-disk-floor.sh 自身的默认值),运维可独立调节两个检查点;inode 下限保持共用。被第一道门放行、但剩余空间不足“8 GiB 加构建自身消耗”的主机,现在能完成构建、通过第二道门、跑完两个 vitest 阶段,而不是在已经付过代价的构建之后死掉。
  • R4-2(悬空镜像裁剪注释的理由经实测为假)——按实测事实重写:本通道对构建传 --no-prune,且 until 对镜像按创建时间筛选,因此该行回收创建超过一天的悬空镜像,当天悬空的镜像等每日 02:30 UTC 清理。仅改注释;没有测试钉注释文本,套件保持全绿基线。

转入后续跟进队列(已验证属实,修复在本 PR 范围之外)

  • R4-1(E2E 通道在同一把构建互斥锁下运行同款带标签裁剪且无时间上限)——已对照 run-e2e-tests.sh 验证。给它加上限意味着改动每个 PR 的 E2E 分片都在运行的共享 CI 脚本;已与建议补在 e2e-workflow.test.js 的钉住测试一并记录进延期发现 issue。按该发现自带的兜底方案,本脚本现在带有一段注释记录这种不对称,避免后来的读者把缺失当作疏漏。线程已回复,保持开放。
  • R4-3(release.yml 的 integration_docker 没有磁盘采样器)——已验证:ci.yml 有六处 DFSAMPLE 循环,release.yml 一处没有。接入 ci: gate heavy jobs on a disk floor and persist pressure samples #10394 确立的采样器需要向本 PR 从未改动的工作流添加步骤;已连完整设计一起记录进延期发现 issue。次要的“两个 vitest 调用之间再加一次门采样”未采纳:点采样无法像流式采样那样在 worker 崩溃后留存,且每加一道门都增加抖动面。线程已回复,保持开放。

需要维护者处理

  • R3-1(PR 描述仍把 DISK_FLOOR_MIN_FREE_KB 写成覆盖变量,并称缓存镜像运行不受影响)——当前 head 上仍然过期,且本轮改动进一步拉开了差距(新增 vitest 前下限旋钮)。机器人无法编辑 PR 描述(工作流只在创建时写入),因此线程回复中附上了中英文两半的精确替换句,供维护者粘贴。线程已回复,保持开放。

对早前延期项的影响

第 2 轮的延期项(fd-7 构建互斥锁内的聚合裁剪时长预算)随之缩小:锁内只剩带标签裁剪和门的带界 docker info。第 3 轮的延期项(新测试共用的裁剪失败 stub 开关)未改动。

验证

  • bash -n .github/scripts/run-release-docker-integration.sh —— 通过
  • npx vitest run scripts/tests/release-workflow.test.js —— 基线 85 通过;修复后 87 通过(新增两个测试:vitest 前下限旋钮、介于两道下限之间的主机放行用例)
  • 变异探针(逐个施加、对聚焦套件运行、随后还原并重跑至全绿):
    • 悬空镜像裁剪改为仅在镜像缺失时运行 → 2 个缓存路径测试失败(R4-4 佐证)
    • BuildKit 缓存裁剪改为仅在镜像缺失时运行 → 同样 2 个测试失败(R4-4 佐证)
    • vitest 前下限改回共用的 8 GiB 默认值 → 7 个下限测试失败(R4-5 佐证)
    • vitest 前下限硬编码(去掉环境变量读取)→ 新增的 vitest 前旋钮测试失败(旋钮佐证)
  • npm run build —— 通过
  • npm run typecheck —— 通过
  • npm run lint —— 通过(pre-commit 钩子的 prettier/eslint 对两个文件也全部通过)
  • npx prettier --check .github/scripts/run-release-docker-integration.sh scripts/tests/release-workflow.test.js —— 通过
  • 集成测试未运行:本次改动是发布通道的 shell 脚本,由 stub 装置套件覆盖;打包后的 CLI 不涉及该路径
  • npm run generate:settings-schema 未运行:未改动任何 settings 源

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。

🧵 Resolved all 3 selected review thread(s). · 已关闭全部选中的 3 条评审线程。

Re-review when you have a moment. After round 10 this bot stops and leaves the PR for a human. · 有空请复审;第 10 轮后本 bot 停止并将 PR 交给人工。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.25.0

@qwen-code-review-bot qwen-code-review-bot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Partially reviewed — gaps disclosed. Suggestions are inline.

1 Suggestion-level finding(s) this review confirmed are already reported on this PR and are not repeated:

  • shared prune-failure stub switch in the two new prune-failure tests — already reported (round-3 deferral at scripts/tests/release-workflow.test.js:2614)

Not reviewed: build-and-test — Test (macos-latest, Node 22.x) and Test (windows-latest, Node 22.x) were skipped in CI and the changed suite ran on Linux only (87 passed), so the new spawn-based cases were never executed on Darwin or Windows.

Not reviewed: closing-issue enumeration — gh in this environment is older than 2.72.0, so the platform's closing-issue references could not be resolved and a second linked issue cannot be ruled out; issue fidelity was evaluated against #13479, which the PR body names explicitly.

Not reviewed: the executable-script lint — .github/scripts/run-release-docker-integration.sh: shellcheck is not installed.

Deferred under the convergence posture (round 5, not a blocker) — recorded, not requested in this round:

  • scripts/tests/release-workflow.test.js:2433 — [probe] no assertion pins the post-build gate's position outside the fd-8 coordinator hold; moving .sh:140 above the unlock leaves all 87 tests green

Convergence: round 5 posted 6 inline comment(s), 4 of them reported for the first time; the previous round posted 6 (5 new). Findings keep coming back to the same files: .github/scripts/run-release-docker-integration.sh (findings in rounds 3, 4; 3 more now). A cluster that keeps producing siblings usually means the fixes are treating instances of a shared root cause — triaging that cause before the next round, or splitting an independent cluster into its own pull request, tends to end the loop faster than fixing them one at a time. (Observation only — nothing was withheld from this review because of this observation.)

Mechanism health: this round did not close cleanly, so it withholds the incremental anchor — and the round it recovered had no anchor this round could use either — none at all, one with no certifier, one certified by an identity other than the one this round runs under, or one this round's fetch refused or resolved to the head — so the next review re-reads the whole diff unless recovery grafts an earlier own anchor that the round running it can use onto the complete work list this round leaves behind, and keeps doing so until a round's marker carries an anchor again or a graft lands that the round running it can use. (Stated, not acted on — this changes nothing about what the round posts.)

中文说明

仅完成部分审查,审查缺口已披露。 建议见行内评论。

本轮确认的 1 条建议级发现已在 PR 上报告过,不再重复发布(列表见上方英文部分)。

未审查(原文为英文):build-and-test — Test (macos-latest, Node 22.x) and Test (windows-latest, Node 22.x) were skipped in CI and the changed suite ran on Linux only (87 passed), so the new spawn-based cases were never executed on Darwin or Windows.

未审查(原文为英文):closing-issue enumeration — gh in this environment is older than 2.72.0, so the platform's closing-issue references could not be resolved and a second linked issue cannot be ruled out; issue fidelity was evaluated against #13479, which the PR body names explicitly.

未审查(原文为英文):the executable-script lint — .github/scripts/run-release-docker-integration.sh: shellcheck is not installed.

收敛姿态下延后(第 5 轮,非阻断)——已记录,本轮不要求修改:共 1 条(原文未翻译,列表见上方英文部分)。

收敛情况:第 5 轮发布了 6 条行内评论,其中 4 条是首次提出;上一轮发布了 6 条(其中 5 条首次提出)。发现反复回到同一批文件:.github/scripts/run-release-docker-integration.sh(第 3、4 轮已出过发现,本轮又有 3 条)。一个不断再生兄弟发现的簇,通常意味着逐条修复只在处理同一根因的实例——先定位并处理该根因,或把独立的簇拆成单独的 PR,通常比逐条修复更快结束循环。(仅为观察——本轮评审未因此扣留任何内容。)

机制健康:本轮未能干净收尾,因而扣留了增量锚点,而它恢复到的那一轮也没有留下本轮可用的锚点——要么完全没有、要么没有认证者、要么由本轮运行身份之外的身份认证、要么被本轮的获取拒绝或解析为头提交——因此下一次评审将重读整个 diff,除非恢复流程把本轮能使用的更早自有锚点嫁接到本轮留下的完整工作清单上;并会一直如此,直到某一轮的标记重新带上锚点,或落地的嫁接能被运行该轮的评审使用。(仅陈述,不据此行动——这不改变本轮发布的任何内容。)

— qwen3.8-max via Qwen Code /review (v0.25.0)

Comment thread .github/scripts/run-release-docker-integration.sh
Comment thread .github/scripts/run-release-docker-integration.sh Outdated
Comment thread scripts/tests/release-workflow.test.js Outdated
Comment thread .github/scripts/run-release-docker-integration.sh Outdated
@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🕐 Review received — an automatic review of the current head is still running, so this round is held until it lands (a push now would cancel it and discard its work, #8888). Your feedback stays queued for the next eligible round.

中文说明

🕐 已收到评审 —— 当前 head 上仍有一轮自动 review 在运行,本轮暂缓(现在推送会取消该 review 并丢弃其工作,#8888)。反馈保持排队,等待下一次可运行的轮次处理。

…mage inspects (#13479)

Address the round-5 review of #13481: forward empty overrides so
check-disk-floor.sh's calibrated defaults apply instead of copied
literals, validate the DISK_FLOOR_DOCKER_* knobs in the lane so a
malformed value fails fast naming the knob this gate reads, re-label
helper failures with those knobs, and bound both docker image inspect
calls (the presence probe discriminates a 124 timeout from "image
absent") so every daemon call under the shared host locks is
time-bounded. The stub helper now counts invocations, so the duplicated
between-floors case becomes the build-path post-gate trip it was named
for.

Co-authored-by: Qwen-Coder <[email protected]>
@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🤖 Addressed the latest review feedback (round 6/10). What changed, and what I pushed back on: · 已处理最新评审反馈(第 6/10 轮)。改动内容与我反驳保留之处如下:

Autofix round summary — PR #13481 (issue #13479)

Commit: fae808c24f fix(release): derive the docker floor from the helper and bound the image inspects (#13479)

All six findings this round trace to the same integration seam in run-release-docker-integration.sh — the two floor-gate checkpoints and the daemon calls made while the shared host locks are held — so they were fixed as one coherent batch instead of six point patches.

Feedback dispositions

rc:4213016563 (R4-1) — resolved in code

The "every daemon call is time-bounded" comment overclaimed: both docker image inspect calls ran unbounded while fds 8/9 were held. Rather than narrowing the wording, the invariant was made true:

  • The image-id read is now timeout 60 docker image inspect --format '{{.Id}}' ..., which fails closed under set -e (no || true, so an empty id can never reach QWEN_SANDBOX_IMAGE).
  • The presence probe is now bounded too, with the asymmetry the finding called out respected: a 124 (timeout) is discriminated from "image absent" and fails the lane fast, because rebuilding against a wedged daemon would hang with the locks held until the job timeout. Any other non-zero still takes the rebuild branch, unchanged.

Witnesses: text pins for the bound and the 124 discrimination in the existing prune/gate text test, plus a new execution case (inspectTimesOut) that simulates the bound via the stubbed pass-through timeout.

rc:4213016572 (R4-5) — resolved in code

The pre-vitest floor and the inode floor no longer copy check-disk-floor.sh's calibrated defaults as literals. Both are forwarded empty (${DISK_FLOOR_DOCKER_POST_MIN_FREE_KB:-} / ${DISK_FLOOR_DOCKER_MIN_FREE_INODES:-}); the helper's :- fallback resolves its own defaults, so a repo-wide retune now reaches these gates by construction and the existing "charges the calibrated job floor" comments are true. The explicit env-prefix assignment stays, so an ambient DISK_FLOOR_MIN_FREE_KB/DISK_FLOOR_MIN_FREE_INODES still cannot move the docker floors (the ambient-isolation cases pass unchanged). check-disk-floor.sh itself was not touched, and the build-sized 8 GiB pre-build budget remains a lane-local literal because it is not the helper's default.

Witness: the stub helper echoes the floors as received, so the affected assertions now pin floor=unset/inodes=unset; re-hardcoding a literal at either call site turns them red (probe 1 and 2 below).

rc:4213016593 (R5-1) — resolved in code

The stub helper now counts its invocations and honors a POST_GATE_EXIT override for calls after the first, driven by a new postGateExit harness option. gateExit: '1' still fails the first checkpoint (the two existing trip cases are untouched and pass). The duplicated between-floors case was converted into the scenario its name claimed — postGateExit: '1' asserting the build ran, the build lock was released, and npx vitest never ran — which is the build-path post-gate trip coverage the mutation battery showed nothing else executed. Its previous assertions survive verbatim in the ordering test under byte-identical input; recorded in test-weakening.json with that evidence.

rc:4213016603 (R5-2) — resolved in code

  • The lane validates all three DISK_FLOOR_DOCKER_* knobs at the top of the gate function (so the pre-build call already validates the pre-vitest knob): a malformed value now fails fast, before the image build, with ::error::<docker knob> must be a non-negative integer of at most 18 digits naming the knob the operator actually set. Empty stays valid (helper default applies). The helper's own validation is untouched.
  • Any helper failure is additionally re-labelled by the lane with the gate's own knobs and the docker data root, so a genuine floor breach no longer reads as the job-start gate's variables.

Witnesses: a new case (dockerPostFloorKb: '3G') asserts non-zero exit, the docker knob named in the output, and no build started; the existing pre-build trip case now also pins the re-labelled failure message.

rc:4213016800 (R3-1) — escalated (maintainer action), thread left open

The remaining gap is the PR body text, which only a maintainer can edit (the workflow writes it at creation). The corrected wording was already drafted in comment 4211178811; a reply on the thread restates that and leaves the thread unresolved as the standing reminder. No code change exists for this finding.

rc:4213016977 (R4-3) — deferred to the follow-up queue, thread left open

The df-sampler instrumentation for integration_docker is verified real, but the remedy adds steps to .github/workflows/release.yml, a file this PR never touches — outside the PR footprint boundary. Recorded in deferred-findings.json with the full design (10-second df loop over the data root and RUNNER_TEMP started before the first vitest, stopped from the existing EXIT trap, samples echoed to the job log as well as the file per the ci.yml:780 pattern, plus an if: failure() upload-artifact step, pinned by a new suite case) so it survives the merge; a thread reply says the same.

Conflict notes

None — --conflict false; no merge with the base was needed.

Verification

  • bash -n .github/scripts/run-release-docker-integration.sh — passed
  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/release-workflow.test.js — 89 passed (87 pre-round + 2 new; the duplicated case was converted, not added)
  • Mutation probes (each: mutate the script, run the suite, confirm red, restore, re-run green 89/89):
    • re-hardcode :-2097152 at the post-gate call site → red (floor=unset assertions fail)
    • re-hardcode :-100000 in the inode forward → red (inodes=unset assertions fail)
    • drop timeout 60 from the image-id read → red (text pin)
    • drop the 124 discrimination from the presence probe → red (timeout case + text pin)
    • remove the knob-validation loop → red (malformed-knob case)
    • revert the failure re-labelling wrapper → red (pre-build trip case's new assertions)
    • make the post-build gate non-fatal (|| true) → red (converted post-gate trip case)
  • node --test .github/scripts/check-disk-floor.test.mjs — 0 failures
  • npm run build — passed
  • npm run typecheck — passed
  • npm run lint — passed
  • npx prettier --check scripts/tests/release-workflow.test.js — passed (and the pre-commit hook's prettier/eslint pass made no changes)

Environment note: this sandbox makes $HOME/.cache read-only, so corepack-backed commands were run with COREPACK_HOME=/tmp/autofix-review-13481/corepack-home. The first unredirected npm run build failed with EACCES: mkdir /home/github-runner/.cache/node/corepack/v1 — an environment constraint, not the change; the redirected rerun passed. No settings source changed; the build's schema regeneration produced no diff.

中文说明

Autofix 本轮总结 — PR #13481(issue #13479)

提交:fae808c24f fix(release): derive the docker floor from the helper and bound the image inspects (#13479)

本轮六条发现都可追溯到 run-release-docker-integration.sh 中的同一处集成接缝——两个磁盘下限门检查点,以及在持有共享主机锁期间发出的守护进程调用——因此作为一批连贯的修改一并处理,而不是六个孤立的点状补丁。

反馈处置

rc:4213016563(R4-1)— 已在代码中解决

“每一次守护进程调用都有时间上限”的注释言过其实:两处 docker image inspect 在持有 fd 8/9 的情况下均无上限运行。本轮没有收窄措辞,而是让该不变量真正成立:

  • 镜像 id 读取改为 timeout 60 docker image inspect --format '{{.Id}}' ...,在 set -e 下失败即关闭(不加 || true,空 id 永远不会进入 QWEN_SANDBOX_IMAGE)。
  • 存在性探测同样加上上限,并遵守发现中指出的不对称约束:124(超时)与“镜像不存在”被显式区分并快速失败,因为对卡死的守护进程重建会在持锁状态下挂到作业超时。其他非零状态仍走重建分支,行为不变。

证据:在现有的裁剪/门文本钉法用例中新增对上限与 124 区分的文本钉,外加一个新的执行用例(inspectTimesOut),通过 stub 的透传 timeout 模拟超时。

rc:4213016572(R4-5)— 已在代码中解决

vitest 前下限与 inode 下限不再以字面值复制 check-disk-floor.sh 的标定默认值。两者都以空值转发(${DISK_FLOOR_DOCKER_POST_MIN_FREE_KB:-} / ${DISK_FLOOR_DOCKER_MIN_FREE_INODES:-});helper 的 :- 回落解析自己的默认值,因此仓库级重新标定现在天然传导到这些门,现有“收取标定作业下限”的注释也真正成立。显式的环境前缀赋值保留,因此环境里的 DISK_FLOOR_MIN_FREE_KB/DISK_FLOOR_MIN_FREE_INODES 仍然无法移动 docker 下限(环境隔离用例原样通过)。check-disk-floor.sh 本身未改动;构建体量的 8 GiB 构建前预算保留为本通道字面值,因为它不是 helper 的默认值。

证据:stub helper 按收到的值原样回显下限,因此受影响的断言现在钉住 floor=unset/inodes=unset;在任何调用点重新硬编码字面值都会让它们变红(见下方探针 1、2)。

rc:4213016593(R5-1)— 已在代码中解决

stub helper 现在统计自身被调用次数,并为第一次之后的调用支持 POST_GATE_EXIT 覆盖,由新的 postGateExit 装置参数驱动。gateExit: '1' 仍然使第一个检查点失败(两个既有触发用例未改动且通过)。重复的“介于两个下限之间”用例被转换为它名字所声称的场景——postGateExit: '1',断言构建已运行、构建锁已释放、npx vitest 从未运行——这正是变异电池显示此前没有任何用例执行到的构建路径 vitest 前门触发覆盖。它原有的断言在相同输入下逐字存留于顺序测试中;已连同该证据记入 test-weakening.json。

rc:4213016603(R5-2)— 已在代码中解决

  • 通道在门函数开头校验全部三个 DISK_FLOOR_DOCKER_* 旋钮(因此构建前调用就已校验 vitest 前旋钮):格式错误的值现在快速失败——在镜像构建之前——并以 ::error::<docker 旋钮> must be a non-negative integer of at most 18 digits 点名运维实际设置的那个旋钮。空值仍然合法(适用 helper 默认值)。helper 自身的校验未改动。
  • helper 的任何失败还会被通道用本门的旋钮与 docker 数据根目录重新标记,因此真实的下限触发不再被读作作业启动门的变量。

证据:新用例(dockerPostFloorKb: '3G')断言非零退出、输出中点名 docker 旋钮、且构建未开始;既有构建前触发用例现在也钉住了重新标记的失败信息。

rc:4213016800(R3-1)— 升级处理(需维护者操作),线程保持开放

剩余缺口是 PR 描述文本,只有维护者能编辑(工作流只在创建时写入)。更正措辞已在评论 4211178811 中起草;线程上的回复重申了这一点,并保持线程未解决,作为持续提醒。该发现不存在代码改动。

rc:4213016977(R4-3)— 延后至跟进队列,线程保持开放

integration_docker 的 df 采样指标缺口已核实为真,但补救需要向 .github/workflows/release.yml 添加步骤,而本 PR 从未改动该文件——超出 PR 足迹边界。已连同完整设计记入 deferred-findings.json(在第一次 vitest 前启动针对数据根目录与 RUNNER_TEMP 的 10 秒 df 循环,由现有 EXIT trap 停止,按 ci.yml:780 的做法把每条采样同时写入作业日志与文件,外加 if: failure() 的 upload-artifact 步骤,并用新的套件用例钉住),确保合并后依然保留;线程回复说了同样的话。

冲突说明

无——--conflict false;无需与基线分支合并。

验证

  • bash -n .github/scripts/run-release-docker-integration.sh — 通过
  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/release-workflow.test.js — 89 通过(上轮 87 + 新增 2;重复用例是转换而非新增)
  • 变异探针(每个均为:改脚本、跑套件、确认变红、还原、复跑全绿 89/89):
    • 在 vitest 前调用点重新硬编码 :-2097152 → 变红(floor=unset 断言失败)
    • 在 inode 转发中重新硬编码 :-100000 → 变红(inodes=unset 断言失败)
    • 移除镜像 id 读取的 timeout 60 → 变红(文本钉)
    • 移除存在性探测的 124 区分 → 变红(超时用例 + 文本钉)
    • 移除旋钮校验循环 → 变红(格式错误旋钮用例)
    • 还原失败重标记包装 → 变红(构建前触发用例的新断言)
    • 让 vitest 前门非致命(|| true)→ 变红(转换后的 vitest 前门触发用例)
  • node --test .github/scripts/check-disk-floor.test.mjs — 0 失败
  • npm run build — 通过
  • npm run typecheck — 通过
  • npm run lint — 通过
  • npx prettier --check scripts/tests/release-workflow.test.js — 通过(pre-commit 钩子的 prettier/eslint 检查未产生改动)

环境说明:本沙箱使 $HOME/.cache 只读,因此依赖 corepack 的命令均以 COREPACK_HOME=/tmp/autofix-review-13481/corepack-home 运行。第一次未重定向的 npm run build 以 EACCES: mkdir /home/github-runner/.cache/node/corepack/v1 失败——这是环境约束而非本次改动;重定向后重跑通过。未改动任何设置来源;构建时的 schema 再生成没有产生 diff。

Deferred non-Critical feedback

Critical-only mode is active: 5 change-producing rounds are complete. The workflow excluded the non-Critical feedback below from this round's actionable sections; the items remain open for human follow-up. Maintainer feedback is deferred only after its author has used 2 regular feedback batches in this window's Critical-only tail; authors at that budget, if any, are named below. (@qwen-code /retry starts a fresh counting window.)

中文说明

已进入仅处理 Critical 的模式:已完成 5 个产生改动的轮次。本轮可执行区域已排除下方非 Critical 反馈;这些条目保持开放,留待人工跟进。维护者反馈仅在其本人于本窗口 Critical-only 阶段已使用 2 批常规反馈预算后才会延后;达到预算的作者(如有)在下方点名。(评论 @qwen-code /retry 可开启新的计数窗口。)

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。

🧵 Resolved all 4 selected review thread(s). · 已关闭全部选中的 4 条评审线程。

Re-review when you have a moment. After round 10 this bot stops and leaves the PR for a human. · 有空请复审;第 10 轮后本 bot 停止并将 PR 交给人工。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.25.0

@qwen-code-review-bot qwen-code-review-bot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Downgraded from Approve to Comment: CI failing: Test (ubuntu-latest, Node 22.x). Partially reviewed — gaps disclosed.

Not reviewed: closing-issue enumeration — gh in this environment is 2.45.0, older than the 2.72.0 the platform's closing-issue references need, so a second linked issue cannot be ruled out; issue fidelity was evaluated against #13479, which the PR body names explicitly, over its fetched body and 22 comments.

Not reviewed: build-and-test — Test (macos-latest, Node 22.x) and Test (windows-latest, Node 22.x) were skipped in CI (neither triggers on pull_request) and their suites did not run locally, so the new spawn-based cases were executed on Linux only; and Test (ubuntu-latest, Node 22.x) failed before reaching test:scripts, so the changed suite has no CI lane evidence at this head (it was run locally instead: 89/89 on the changed file, 2864 passed across the full 104-file scripts suite).

Not explored to full depth (tool budget reached): "agent 1c": could not execute shellcheck --enable=all --severity=style on .github/scripts/run-release-docker-integration.sh (binary absent from the worktree, downloaded…; "agent 6c": executing the 16 new spawnSync cases — the review worktree has no node_modules , and installing into a tree shared with the other review agents would leave s….

Deferred under the convergence posture (round 6, not a blocker) — recorded, not requested in this round:

  • .github/scripts/run-release-docker-integration.sh:49 — [probe] two of the three knob names in the new validation loop are deletable with the suite green (carried from R5-2, fix-induced)
  • .github/scripts/run-release-docker-integration.sh:59 — [probe] a 124 from docker info silently disables the gate while the same 124 twelve lines below is fatal, and the daemon's diagnostic is discarded
  • .github/scripts/run-release-docker-integration.sh:74 — [probe] the empty-forward coupling with the helper's :- default is pinned by nothing on either side (carried from R4-5, fix-induced)
  • .github/scripts/run-release-docker-integration.sh:152 — [probe] the bound added at R4-1's request is pinned by text only, so its fail-closed behaviour and exported value are untested (carried from R4-1, fix-induced)
  • .github/scripts/run-release-docker-integration.sh:152 — [probe] the newly bounded id read is the only failure path this diff touches that dies with no GitHub annotation

Mechanism health: this round did not close cleanly, so it withholds the incremental anchor — and the round it recovered had no anchor this round could use either — none at all, one with no certifier, one certified by an identity other than the one this round runs under, or one this round's fetch refused or resolved to the head — so the next review re-reads the whole diff unless recovery grafts an earlier own anchor that the round running it can use onto the complete work list this round leaves behind, and keeps doing so until a round's marker carries an anchor again or a graft lands that the round running it can use. (Stated, not acted on — this changes nothing about what the round posts.)

中文说明

⚠️ 已从批准降级为评论:CI failing: Test (ubuntu-latest, Node 22.x)。 仅完成部分审查,审查缺口已披露。

未审查(原文为英文):closing-issue enumeration — gh in this environment is 2.45.0, older than the 2.72.0 the platform's closing-issue references need, so a second linked issue cannot be ruled out; issue fidelity was evaluated against #13479, which the PR body names explicitly, over its fetched body and 22 comments.

未审查(原文为英文):build-and-test — Test (macos-latest, Node 22.x) and Test (windows-latest, Node 22.x) were skipped in CI (neither triggers on pull_request) and their suites did not run locally, so the new spawn-based cases were executed on Linux only; and Test (ubuntu-latest, Node 22.x) failed before reaching test:scripts, so the changed suite has no CI lane evidence at this head (it was run locally instead: 89/89 on the changed file, 2864 passed across the full 104-file scripts suite).

未探索到全部深度(达到工具调用预算):"agent 1c":could not execute shellcheck --enable=all --severity=style on .github/scripts/run-release-docker-integration.sh (binary absent from the worktree, downloaded…;"agent 6c":executing the 16 new spawnSync cases — the review worktree has no node_modules , and installing into a tree shared with the other review agents would leave s…。

收敛姿态下延后(第 6 轮,非阻断)——已记录,本轮不要求修改:共 5 条(原文未翻译,列表见上方英文部分)。

机制健康:本轮未能干净收尾,因而扣留了增量锚点,而它恢复到的那一轮也没有留下本轮可用的锚点——要么完全没有、要么没有认证者、要么由本轮运行身份之外的身份认证、要么被本轮的获取拒绝或解析为头提交——因此下一次评审将重读整个 diff,除非恢复流程把本轮能使用的更早自有锚点嫁接到本轮留下的完整工作清单上;并会一直如此,直到某一轮的标记重新带上锚点,或落地的嫁接能被运行该轮的评审使用。(仅陈述,不据此行动——这不改变本轮发布的任何内容。)

— qwen3.8-max via Qwen Code /review (v0.25.0)

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🕐 Review received — an automatic review of the current head is still running, so this round is held until it lands (a push now would cancel it and discard its work, #8888). Your feedback stays queued for the next eligible round.

中文说明

🕐 已收到评审 —— 当前 head 上仍有一轮自动 review 在运行,本轮暂缓(现在推送会取消该 review 并丢弃其工作,#8888)。反馈保持排队,等待下一次可运行的轮次处理。

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🔀 Base updated: red check(s) [Test (ubuntu-latest, Node 22.x)] pass on current main — merged current main via update-branch; CI will re-run.

中文说明

🔀 已更新 base:红色检查 [Test (ubuntu-latest, Node 22.x)] 在当前 main 上通过 —— 已通过 update-branch 合入当前 main,CI 将重新运行。

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Release Failed for v0.25.0-nightly.20261005.69d5db2ff2 on 2026-10-05

4 participants