Skip to content

fix(ci): tolerate unwritable docker sandbox lock dir on self-hosted runners (#12006) - #12016

Closed
qwen-code-dev-bot wants to merge 14 commits into
mainfrom
autofix/issue-12006
Closed

qwen-code-dev-bot wants to merge 14 commits into
mainfrom
autofix/issue-12006

Conversation

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator

What this PR does

The self-hosted Linux docker E2E leg opens its host-side flock files (daemon prune-exclusion, per-commit build coordinator, build mutex) under ${HOME}/.cache/qwen-code-ci. When a root-owned leftover makes that directory unusable, every exec 9>-style open dies with EACCES before a single test runs. This PR routes both lock-opening call sites — the E2E runner script and the workflow's image-prune step — through a small resolver script that prints the shared directory when it is genuinely usable (directory writable and every existing lock file writable) and otherwise prints a job-private fallback under the job's temp dir, with a ::warning:: in the log. On healthy hosts the resolved path is the same shared directory as before, so the locking protocol is unchanged; on a degraded host the leg now runs with per-job coordination instead of dying in under a second.

Why it's needed

Main-branch E2E run 35069321648 (tracked in #12006) failed the sandbox:docker leg's Run E2E tests and Prune dangling docker images steps in under one second each. The job's Restore workspace ownership step had already warned could not restore CI cache ownership: its chown … || sudo -n chown … chain cannot repair a root-owned lock directory on a runner whose user has no passwordless sudo (a non-root process can never chown another user's files). This is the same failure class as #11990 recurring on a host where the heal is powerless, so the workflow must tolerate the unwritable directory instead of depending on healing it.

Losing cross-job coordination on such a host is safe: the sandbox image build is idempotent (a duplicate concurrent build wastes work but converges on the same tag), and both prune paths are filtered by label and until=24h, so they cannot remove an image a running leg is using — the host cleanup service already relies on the age gate rather than the lock for exactly this reason.

Reviewer Test Plan

How to verify

The change is CI-runner shell behavior; the review surface is the new resolver and its two call sites.

  1. Confirm the failure mode being fixed: in run 35069321648 both failed steps lasted <1s and the only warning was the failed cache-ownership heal — i.e. the lock open died before tests ran. The previous incident (Main CI failed: E2E Tests on dfad77be8c66 #11990) documented the EACCES mechanism.
  2. Check the resolver contract: .github/scripts/resolve-ci-lock-dir.sh prints the shared dir only when it is writable along with every existing *.lock, otherwise falls back to ${RUNNER_TEMP}/qwen-code-ci-locks and warns on stderr (stdout stays a clean single path for $(...) capture).
  3. Run the witnesses: npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/resolve-ci-lock-dir.test.js scripts/tests/e2e-workflow.test.js — the new suite executes the real runner script's docker branch against an unwritable lock file (the exact Main CI failed: E2E Tests on aff26a8f36df #12006 reproduction: exit 1 EACCES pre-fix, green post-fix) and against a healthy home; the workflow suite pins both call sites to the resolver.
  4. Expected behavior on the next main E2E run: a host with a poisoned lock dir logs one ::warning:: about the job-private fallback and the docker leg runs normally; a healthy host is byte-for-byte the old behavior.

Evidence (Before & After)

N/A — CI infrastructure only. Before: run-e2e-tests.sh: line 46: …/docker-sandbox-daemon.lock: Permission denied, step exit 1 in <1s (reproduced locally with a chmod 0400 lock file). After: the same poisoned setup completes the leg on a job-private lock dir with a warning.

Tested on

OS Status
🍏 macOS ⚠️ not tested
🪟 Windows N/A (bash-driven suite excluded there, matching the existing suites)
🐧 Linux ✅ tested

Environment (optional)

Autofix container (no docker daemon): bash-level execution of the real runner script with stubbed docker/npx, plus the repository's workflow-contract vitest suites. The real docker leg cannot run here; the next main-branch E2E run is the end-to-end confirmation.

Risk & Scope

  • Main risk or tradeoff: on a host with an unwritable shared lock dir, concurrent jobs lose build/prune coordination (duplicate image builds are possible; prune could run during another leg). Both are benign — builds are idempotent and prunes are label- plus age-filtered — and the alternative is the leg never running at all.
  • Not validated / out of scope: the real sandbox:docker E2E leg (no docker daemon in the authoring environment); yamllint (no pip in the container — the two-line YAML edit preserves indentation and actionlint parses the file cleanly). .github/scripts/run-release-docker-integration.sh and sdk-java.yml use the same shared lock dir and keep the same exposure; fixing them is a deliberate follow-up to keep this PR scoped to the failing workflow. A pre-existing verify-capture.test.js ANSI-rendering failure exists in this container and reproduces identically on clean HEAD; it is unrelated.
  • Breaking changes / migration notes: none. The Restore workspace ownership heal step is kept as-is — where sudo exists it still restores full shared-dir coordination.

Linked Issues

Fixes #12006

中文说明

本次 PR 的内容

自托管 Linux docker E2E 分支在 ${HOME}/.cache/qwen-code-ci 下打开宿主机 flock 文件(daemon prune 互斥、按提交构建协调、构建互斥)。当 root 属主的残留导致该目录不可用时,每个 exec 9> 式打开都会以 EACCES 失败,连一个测试都来不及运行。本 PR 将两个打开锁的调用点——E2E 运行脚本和工作流中的镜像清理步骤——改为通过一个小型解析脚本获取锁目录:共享目录确实可用(目录可写且所有现存锁文件可写)时输出它,否则输出作业临时目录下的作业私有回退目录,并在日志中打印 ::warning::。主机健康时解析结果与原先完全相同,锁定协议不变;主机受损时该分支以作业内协调运行,而不是在一秒内死亡。

为什么需要

main 分支 E2E 运行 35069321648(由 #12006 跟踪)中,sandbox:docker 分支的 Run E2E tests 和 Prune dangling docker images 两个步骤都在不到一秒内失败。该作业的 Restore workspace ownership 步骤已经告警 could not restore CI cache ownership:其 chown … || sudo -n chown … 链在 runner 用户没有免密 sudo 时无法修复 root 属主的锁目录(非 root 进程永远无法 chown 他人的文件)。这与 #11990 是同一故障类别在修复手段失效的主机上的复发,因此工作流必须容忍目录不可写,而不是依赖修复它。

在这类主机上失去跨作业协调是安全的:sandbox 镜像构建是幂等的(并发的重复构建只是浪费算力,最终收敛到同一个 tag),而两条 prune 路径都带标签和 until=24h 过滤,不可能删除正在运行的分支在用的镜像——宿主机清理服务本来就依赖时限而非锁来保证这一点。

审查者测试计划

如何验证

本次改动是 CI runner 的 shell 行为;审查重点是新的解析脚本及其两个调用点。

  1. 确认被修复的故障模式:运行 35069321648 中两个失败步骤耗时 <1 秒,唯一警告是缓存属主修复失败——即锁的打开在测试运行前就已死亡。此前的事故(Main CI failed: E2E Tests on dfad77be8c66 #11990)记录了 EACCES 机理。
  2. 检查解析器契约:.github/scripts/resolve-ci-lock-dir.sh 仅在共享目录可写且所有现存 *.lock 可写时输出它,否则回退到 ${RUNNER_TEMP}/qwen-code-ci-locks 并向 stderr 告警(stdout 保持为干净的单一路径供 $(...) 捕获)。
  3. 运行见证测试:npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/resolve-ci-lock-dir.test.js scripts/tests/e2e-workflow.test.js——新套件用不可写锁文件真实执行 runner 脚本的 docker 分支(即 Main CI failed: E2E Tests on aff26a8f36df #12006 的精确复现:修复前退出码 1 EACCES,修复后转绿)以及健康 home 的情形;工作流套件锚定两个调用点都经过该解析器。
  4. 下一次 main E2E 运行的预期行为:锁目录中毒的主机记录一条关于作业私有回退的 ::warning:: 并正常跑完 docker 分支;健康主机与旧行为逐字节一致。

前后对比证据

N/A——纯属 CI 基础设施。修复前:run-e2e-tests.sh: line 46: …/docker-sandbox-daemon.lock: Permission denied,步骤 <1 秒退出码 1(已用 chmod 0400 锁文件在本地复现)。修复后:同样的中毒设置在作业私有锁目录上完整跑完该分支并输出一条警告。

已测试平台

操作系统 状态
🍏 macOS ⚠️ 未测试
🪟 Windows N/A(bash 驱动套件在该平台排除,与既有套件一致)
🐧 Linux ✅ 已测试

环境(可选)

autofix 容器(无 docker 守护进程):在 stub 的 docker/npx 下真实执行 runner 脚本的 bash 级验证,外加仓库的工作流契约 vitest 套件。真实的 docker 分支无法在此运行;下一次 main 分支 E2E 运行是端到端确认。

风险与范围

  • 主要风险或取舍:在共享锁目录不可写的主机上,并发作业失去构建/清理协调(可能出现重复镜像构建;清理可能在其他分支运行期间执行)。两者都无害——构建幂等,清理有标签和时限过滤——而备选结局是该分支根本跑不起来。
  • 未验证 / 超出范围:真实的 sandbox:docker E2E 分支(编写环境无 docker 守护进程);yamllint(容器内无 pip——两行 YAML 改动保持缩进一致,actionlint 已干净解析)。.github/scripts/run-release-docker-integration.sh 与 sdk-java.yml 使用同一共享锁目录、保留同样的暴露面;修复它们是有意留下的后续项,以保持本 PR 只覆盖出故障的工作流。容器内还存在一个既有的 verify-capture.test.js ANSI 渲染失败,在干净 HEAD 上可同样复现,与本次改动无关。
  • 破坏性变更 / 迁移说明:无。Restore workspace ownership 修复步骤保持原样——在有 sudo 的环境中它仍能恢复完整的共享目录协调。

关联 Issue

Fixes #12006

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

Autofix E2E Report — issue #12006

Issue

Main-branch E2E run 35069321648 on aff26a8f36 failed the E2E Test (Linux) - sandbox:docker - shard 1/1 job in the steps Run E2E tests and Prune dangling docker images.

Diagnosis (evidence)

  • The public jobs API shows both failed steps completed in under one second (07:46:47 → 07:46:47) — no test, build, or prune actually ran. The sibling sandbox:none leg and both macOS legs passed, so the commit's own E2E test changes were not the trigger.
  • The job's only annotation besides the two exit code 1 failures was the Restore workspace ownership step warning: could not restore CI cache ownership; docker sandbox locks may stay unreadable. That heal (added by fix(ci): heal the docker lock cache ownership in e2e-test-linux (#11990) #11992 for Main CI failed: E2E Tests on dfad77be8c66 #11990) chains chown … || sudo -n chown …; a non-root runner can never chown a root-owned file, and this pool runner has no passwordless sudo, so the heal warned and gave up.
  • Both failed steps open the same host-side flock file with a write redirection under set -e/bash -e: exec 9>"${HOME}/.cache/qwen-code-ci/docker-sandbox-daemon.lock" (in .github/scripts/run-e2e-tests.sh for the test step, and in e2e.yml for the prune step). A root-owned leftover there fails the open with EACCES and kills the step before any work — exactly the observed 0-second deaths.

Reproduction

Docker is not available in this autofix container, so the exact CI leg cannot run here. Instead the failure was reproduced at bash level by executing the real run-e2e-tests.sh (docker branch, RUNNER_ENVIRONMENT=self-hosted, stubbed docker/npx) against a $HOME/.cache/qwen-code-ci/docker-sandbox-daemon.lock made unwritable (chmod 0400):

  • Pre-fix script: run-e2e-tests.sh: line 46: …/docker-sandbox-daemon.lock: Permission denied, exit 1 — the same instant death as the CI steps.
  • Post-fix script: the leg runs to completion (exit 0) on a job-private lock dir, emitting a ::warning:: about the degraded coordination.

Fix

Both lock-opening call sites now resolve the lock directory through a new helper, .github/scripts/resolve-ci-lock-dir.sh, instead of hardcoding ${HOME}/.cache/qwen-code-ci. The helper prints the shared dir when it is usable (mkdir succeeds, dir writable, every existing *.lock writable); otherwise it prints a job-private fallback under $RUNNER_TEMP and warns. Losing cross-job coordination on a degraded host is safe: the sandbox image build is idempotent, and both prune paths are label- and age-filtered (until=24h), so they never touch an image a running leg could be using. When the host is healthy (or the heal step can chown, e.g. where sudo exists), behavior is byte-for-byte the old shared-dir protocol.

The heal step itself is unchanged — it still restores full coordination where it has the privileges. .github/scripts/run-release-docker-integration.sh and sdk-java.yml open locks in the same shared dir and remain exposed to the same host-state failure; they are separate workflows and were left untouched to keep this fix scoped (noted for follow-up).

Regression coverage

scripts/tests/resolve-ci-lock-dir.test.js (new, bash-executed, excluded on Windows like the other bash-driven suites) witnesses: resolver prefers the shared dir when writable; falls back on an unwritable lock file; falls back on an unwritable dir; the full docker branch of run-e2e-tests.sh completes on the shared dir when healthy; and — the #12006 reproduction — completes via the job-private fallback when the daemon lock is unwritable. A mutation probe confirmed the reproduction test fails against the pre-fix script (exit 1, EACCES) and passes with the fix. scripts/tests/e2e-workflow.test.js pins were updated to the new ${ci_lock_dir} shape and now also pin that both call sites resolve through the helper before opening any descriptor.

Verification

  • Reproduction probe (pre-fix script, unwritable lock file): exit 1 with …/docker-sandbox-daemon.lock: Permission denied — matches the CI failure signature.
  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/resolve-ci-lock-dir.test.js scripts/tests/e2e-workflow.test.js scripts/tests/e2e-shard-retry.test.js — 47 passed.
  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/resolve-ci-lock-dir.test.js scripts/tests/e2e-workflow.test.js scripts/tests/e2e-shard-retry.test.js scripts/tests/release-workflow.test.js scripts/tests/unit-vitest-configs.test.ts — 165 passed.
  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/sdk-java-workflow.test.js scripts/tests/no-ak-integration-ci.test.js — 19 passed.
  • Mutation probe: witness test re-run against the pre-fix run-e2e-tests.sh — fails with exit 1 (EACCES) as expected; restored script passes.
  • bash .github/scripts/check-workflow-size.sh — passed (e2e.yml stays within the 4096-byte growth allowance; the qwen-autofix.yml size warning is pre-existing on main).
  • shellcheck 0.11.0 with the repo's flags (--check-sourced --enable=all --exclude=SC2002,SC2129,SC2310 --severity=style) — new helper is clean; run-e2e-tests.sh shows only the 3 pre-existing warnings, zero new findings.
  • actionlint 1.7.12 with the repo's flags on .github/workflows/e2e.yml — clean.
  • npx prettier --experimental-cli --check on all changed files — passed (after --write normalized the two test files).
  • npx eslint scripts/tests/vitest.config.ts scripts/tests/e2e-workflow.test.js scripts/tests/resolve-ci-lock-dir.test.js — passed.
  • npm run lint — passed.
  • npm run typecheck — passed.
  • npm run build — passed.
  • npm run test:scripts (full scripts suite) — 2494 passed, 17 skipped, 1 failed: scripts/tests/verify-capture.test.js > preserves colour and bold independently. Reproduced identically on clean HEAD with all changes stashed, so it is a pre-existing, environment-specific failure of this container (bold ANSI rendering), unrelated to this change.
  • Not run locally: yamllint (no pip/pip3/uv in this container; the 2-line YAML edit preserves indentation and quoting, and actionlint parsed the document cleanly) and the real docker E2E leg (no docker daemon available; bash-level surrogate described above). The next main-branch E2E run is the end-to-end confirmation.
中文说明

Autofix E2E 报告 — issue #12006

问题

main 分支在 aff26a8f36 上的 E2E 运行 35069321648 中,E2E Test (Linux) - sandbox:docker - shard 1/1 作业在 Run E2E tests 和 Prune dangling docker images 两个步骤失败。

诊断(证据)

  • 公开的 jobs API 显示两个失败步骤都在不到一秒内结束(07:46:47 → 07:46:47)——测试、构建、清理都没有真正运行。同次运行的 sandbox:none 分支和两个 macOS 分支全部通过,因此触发失败的并不是该提交自身的 E2E 测试改动。
  • 该作业除了两条 exit code 1 失败注解外,唯一的注解来自 Restore workspace ownership 步骤的警告:could not restore CI cache ownership; docker sandbox locks may stay unreadable。该修复链(fix(ci): heal the docker lock cache ownership in e2e-test-linux (#11990) #11992 为 Main CI failed: E2E Tests on dfad77be8c66 #11990 加入)是 chown … || sudo -n chown …;非 root 用户永远无法 chown root 拥有的文件,而此池化 runner 没有免密 sudo,因此修复步骤只能告警后放弃。
  • 两个失败步骤都在 set -e/bash -e 下用写重定向打开同一个宿主机 flock 文件:exec 9>"${HOME}/.cache/qwen-code-ci/docker-sandbox-daemon.lock"(测试步骤在 .github/scripts/run-e2e-tests.sh 中,清理步骤在 e2e.yml 中)。该处的 root 属主残留会让打开操作以 EACCES 失败,在真正开始任何工作之前就杀死步骤——与观测到的 0 秒死亡完全一致。

复现

本 autofix 容器中没有 docker,无法运行完全一致的 CI 分支。因此改为在 bash 层面复现:以真实 run-e2e-tests.sh(docker 分支、RUNNER_ENVIRONMENT=self-hosted、stub 的 docker/npx)对一个被改为不可写(chmod 0400)的 $HOME/.cache/qwen-code-ci/docker-sandbox-daemon.lock 执行:

  • 修复前脚本:run-e2e-tests.sh: line 46: …/docker-sandbox-daemon.lock: Permission denied,退出码 1——与 CI 步骤的瞬间死亡一致。
  • 修复后脚本:该分支使用作业私有锁目录完整跑完(退出码 0),并输出一条关于协调降级的 ::warning::。

修复

两个打开锁的调用点不再硬编码 ${HOME}/.cache/qwen-code-ci,改为通过新辅助脚本 .github/scripts/resolve-ci-lock-dir.sh 解析锁目录。辅助脚本在共享目录可用时(mkdir 成功、目录可写、所有现存 *.lock 可写)输出共享目录;否则输出 $RUNNER_TEMP 下的作业私有回退目录并告警。在受损主机上失去跨作业协调是安全的:sandbox 镜像构建是幂等的,两条 prune 路径都带标签和时限过滤(until=24h),绝不会触碰正在运行的分支可能在用的镜像。当主机健康时(或修复步骤有权限 chown 时,例如有 sudo),行为与原先的共享目录协议逐字节一致。

修复步骤本身未改动——在有权限的环境中它仍能恢复完整协调。.github/scripts/run-release-docker-integration.sh 与 sdk-java.yml 也在同一共享目录打开锁、暴露于同样的主机状态故障;它们属于其他工作流,为保持本次修复范围最小而未改动(已记录为后续跟进项)。

回归覆盖

新增的 scripts/tests/resolve-ci-lock-dir.test.js(bash 实际执行,与其他 bash 驱动套件一样在 Windows 通道排除)见证:解析器在共享目录可写时优先使用它;锁文件不可写时回退;目录不可写时回退;run-e2e-tests.sh 的完整 docker 分支在健康时跑在共享目录上;以及——#12006 的复现——在 daemon 锁不可写时通过作业私有回退跑完。变异探针已确认该复现测试在修复前脚本上失败(退出码 1,EACCES),修复后通过。scripts/tests/e2e-workflow.test.js 的锚点已更新为新的 ${ci_lock_dir} 形态,并新增锚点保证两个调用点在打开任何描述符之前都经过该辅助脚本解析。

验证

  • 复现探针(修复前脚本 + 不可写锁文件):退出码 1,报 …/docker-sandbox-daemon.lock: Permission denied——与 CI 失败特征一致。
  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/resolve-ci-lock-dir.test.js scripts/tests/e2e-workflow.test.js scripts/tests/e2e-shard-retry.test.js —— 47 个通过。
  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/resolve-ci-lock-dir.test.js scripts/tests/e2e-workflow.test.js scripts/tests/e2e-shard-retry.test.js scripts/tests/release-workflow.test.js scripts/tests/unit-vitest-configs.test.ts —— 165 个通过。
  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/sdk-java-workflow.test.js scripts/tests/no-ak-integration-ci.test.js —— 19 个通过。
  • 变异探针:将见证测试对修复前的 run-e2e-tests.sh 重跑——按预期失败(退出码 1,EACCES);恢复修复后脚本后通过。
  • bash .github/scripts/check-workflow-size.sh —— 通过(e2e.yml 仍在 4096 字节增长额度内;qwen-autofix.yml 的体积警告是 main 上既有的)。
  • shellcheck 0.11.0(使用仓库参数 --check-sourced --enable=all --exclude=SC2002,SC2129,SC2310 --severity=style)——新辅助脚本干净;run-e2e-tests.sh 仅有 3 条既有警告,无新增问题。
  • actionlint 1.7.12(使用仓库参数)检查 .github/workflows/e2e.yml —— 干净。
  • npx prettier --experimental-cli --check 检查全部改动文件 —— 通过(先以 --write 规整了两个测试文件)。
  • npx eslint scripts/tests/vitest.config.ts scripts/tests/e2e-workflow.test.js scripts/tests/resolve-ci-lock-dir.test.js —— 通过。
  • npm run lint —— 通过。
  • npm run typecheck —— 通过。
  • npm run build —— 通过。
  • npm run test:scripts(完整 scripts 套件)—— 2494 通过、17 跳过、1 失败:scripts/tests/verify-capture.test.js > preserves colour and bold independently。在暂存全部改动的干净 HEAD 上以同样方式复现失败,因此它是本容器环境特有的既有失败(bold ANSI 渲染),与本次改动无关。
  • 本地未运行:yamllint(容器内无 pip/pip3/uv;该 2 行 YAML 改动保持缩进与引用一致,且 actionlint 已干净解析文档)以及真实的 docker E2E 分支(无可用 docker 守护进程;上文描述了 bash 层面的替代验证)。下一次 main 分支 E2E 运行是端到端的最终确认。

🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.23.4

@github-actions github-actions Bot added the review/self-reported The linked issue was opened by the PR author (self-reported) label Sep 16, 2026
@qwen-code-dev-bot

qwen-code-dev-bot commented Sep 16, 2026 •

Copy link
Copy Markdown
Collaborator Author

✅ AutoFix round 3 finished — view run. See this round's report below.

中文说明

✅ AutoFix 第 3 轮已完成 —— 查看运行。本轮报告见下方。

…#12006)

Review round on #12016:

- Stub flock(1) in the docker-leg test harness: it is util-linux and
  absent on macOS, where the new cases would fail the nightly lane.
- The resolver now takes the lock names the caller will open and probes
  only those, so a foreign unwritable lock (sdk-java-tests.lock) no
  longer vetoes a usable shared dir.
- The resolver header no longer claims the fallback is coordination-safe:
  RUNNER_TEMP is per-job, so coordination with sibling jobs, the root
  cleanup timer, and the release lane is lost; the fallback also lands on
  GITHUB_STEP_SUMMARY instead of only in a step log.
- The always()-gated prune step tolerates a missing checkout (resolver
  absent) by falling back to the job-private dir inline.
- Diagnostics name the offending path per probe failure instead of
  asserting an unverified root-owned cause.
- Pin the fail-loud resolver branch and the exec 8> coordinator lock's
  directory, both previously unwitnessed.

Co-authored-by: Qwen-Coder <[email protected]>
@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🤖 Addressed the latest review feedback (round 1/10). What changed, and what I pushed back on: · 已处理最新评审反馈(第 1/10 轮)。改动内容与我反驳保留之处如下:

Autofix review round — PR #12016

Addressed all 7 findings from the latest review (1 Critical + 6 Suggestions). No declines, no deferrals, no open questions.

Findings

R1-1 [Critical] — docker-leg test cases fail where flock is absent (rc:4025497745)

Implemented. Reproduced before fixing: with a PATH farm containing every real binary except flock, the suite failed exactly as reported — 2 failed | 3 passed, both docker-leg cases dying at flock --shared --wait 1800 9 (bash 127 → ::error::docker daemon read lock not acquired within 30 minutes → exit 1). Fix per the suggestion: an exit 0 flock stub in world.bin beside the existing docker/npx stubs — these cases witness lock-DIR resolution, not lock semantics — so the #12006 reproduction stays executed, not skipped, on every lane including test_macos. run-e2e-tests.sh is untouched (the flock --shared --wait 1800 9 pin stands) and nothing was added to the win32 exclusion list.

R1-3 [Suggestion] — a foreign unwritable lock vetoes the shared dir (rc:4025497777)

Implemented. ci_lock_dir_writable now takes the names of the locks the caller will open and probes only those; the directory-level checks stay unconditional, so an unwritable directory still always falls back. run-e2e-tests.sh passes docker-sandbox-daemon.lock docker-sandbox-build.lock "docker-sandbox-build-e2e-${GITHUB_SHA}.lock"; the prune step passes docker-sandbox-daemon.lock; both text pins were updated in the same change. New case: an unwritable runner-owned sdk-java-tests.lock beside a writable daemon lock → the resolver still prints the shared dir (red before, green after). No rm-based healing — lock identity lives on the inode.

R1-2 [Suggestion] — per-job fallback silently drops coordination; the header overclaims (rc:4025497792)

Implemented the "cheapest honest fix" arm. The header now states that coordination with sibling jobs, the root qwen-docker-cleanup timer, and the release lane's build mutex is lost, not merely degraded; that the until=24h age filter does not protect a reused image; and that clearing the state needs a human on the host. When GITHUB_STEP_SUMMARY is set, the resolver also appends the notice there, so the degradation lands on the run page instead of only in a step log. I did not take the host-shared-fallback arm: it adds a second lock namespace to reason about and still cannot restore coordination with the root timer (which keys on the primary path), so it buys little over honesty for a state that should be rare and human-cleared anyway.

R1-4 [Suggestion] — prune step dies at 127 when checkout failed (rc:4025497806)

Implemented. The always()-gated prune step now carries the suggested || { ci_lock_dir="${RUNNER_TEMP:-/tmp}/qwen-code-ci-locks"; mkdir -p ...; } guard, with a comment noting the softening belongs to this step only — the test leg keeps failing loudly. Verified behaviorally in an empty workspace under bash -e: with the guard the step exits 0 and both prunes run; without it the step aborts at 127 and no prune runs.

R1-7 [Suggestion] — the fail-loud branch had no test (rc:4025497840)

Implemented. New case fails loudly when the job-private fallback cannot be created: shared dir at 0555, RUNNER_TEMP under a regular file → asserts exit 1, empty stdout, and ::error:: on stderr. (One deviation from the snippet: the shared dir is made unwritable before the spawn, otherwise the resolver takes the primary branch and never reaches the fallback.)

R1-8 [Suggestion] — the warning asserted an unverified root-owned cause (rc:4025497849)

Implemented. The probe now records why it failed on each exit — cannot create <dir> / <dir> is not writable / lock file <path> is not writable — and the warning and error name the offending path. The hardcoded (unhealed root-owned leftover) is gone. The existing lock-file fallback case now asserts the warning names docker-sandbox-daemon.lock; stdout remains exactly one line — the resolved path — for both callers' $(...) capture.

R1-6 [Suggestion] — exec 8> was not pinned to the resolved dir (rc:4025497864)

Implemented. The bare filename-fragment pin is now the full line exec 8>"${ci_lock_dir}/docker-sandbox-build-e2e-${GITHUB_SHA}.lock" (the commit-qualified name stays), and a new behavioral case poisons the shared directory (0555), runs the docker leg with the image absent, and asserts exit 0 plus all three lock files — daemon, per-commit coordinator, build mutex — created under the job-private fallback. Exercising the build-mutex path required stubbing npm and making the docker stub report the image absent.

Mutation probes

Each run against this round's tree and restored afterwards (restores verified byte-identical):

  • Drop the flock stub and run with the flock-less PATH farm → all 3 docker-leg cases red; the suite is green again with the stub on the same farm (the macOS condition).
  • Remove the resolver's named-lock probe loop → falls back ... when a lock file is not writable and the docker-leg lock-file case red — the narrowing is not a no-op.
  • Revert exec 8> to the ${HOME}/.cache/qwen-code-ci spelling → the new toContain pin AND the directory-poisoning behavioral case both red.
  • Delete the fail-loud exit 1 → the new fail-loud case red (previously green, per the finding).
  • Drop the filename from the per-lock warning → the naming assertion red.
  • Remove the $GITHUB_STEP_SUMMARY block → the step-summary case red.
  • Remove the prune step's || guard → the guard pin red.

Red-before/green-after witnesses (new assertions against the pre-fix scripts): the foreign-lock case, the step-summary case, the warning-names-the-file assertion, and all three updated/new e2e-workflow.test.js pins were red; everything went green after the script changes.

Verification

  • npm run build — passed
  • npm run typecheck — passed
  • npm run lint — passed (eslint; the repo's shellcheck lane is lint:all, which downloads the binary and cannot run in this sandbox — both shell scripts pass bash -n and follow the repo's shellcheck-clean conventions)
  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/resolve-ci-lock-dir.test.js scripts/tests/e2e-workflow.test.js scripts/tests/e2e-shard-retry.test.js scripts/tests/no-ak-integration-ci.test.js — 66 passed
  • node --test .github/scripts/ci-runner-routing.test.mjs — 23 passed
  • npm run test:scripts (full scripts suite) — 2499 passed, 1 failed: verify-capture.test.js > preserves colour and bold independently, a terminal-ANSI rendering assertion reproduced identically on the pre-round tree via a stash probe; it is unrelated to any file this PR touches. (install-script.test.js also failed before npm run build for want of packages/audio-capture/dist in this environment and is green after it.)
  • flock-less PATH farm (/tmp/no-flock-farm: 540 entries = every real PATH binary minus flock) — resolver suite green with the stub, i.e. under the macOS-lane condition.
  • npx prettier --check on all touched files — clean.
中文说明

Autofix 审查轮次 — PR #12016

已处理最新一轮审查的全部 7 条意见(1 条 Critical + 6 条 Suggestion)。无拒绝、无推迟、无待决问题。

意见处理

R1-1 [Critical] — 缺少 flock 时 docker 分支测试用例失败 (rc:4025497745)

已实现。 先复现后修复:用一个包含除 flock 以外全部真实二进制的 PATH 农场运行套件,失败情况与报告完全一致——2 failed | 3 passed,两个 docker 分支用例都死在 flock --shared --wait 1800 9(bash 127 → ::error::docker daemon read lock not acquired within 30 minutes → exit 1)。按建议修复:在 world.bin 中现有的 docker/npx stub 旁边加入一个 exit 0 的 flock stub——这些用例见证的是锁目录的解析,而不是锁语义——因此 #12006 的复现在每条泳道(包括 test_macos)上都保持被执行而非被跳过。run-e2e-tests.sh 未改动(flock --shared --wait 1800 9 锚定保持不变),win32 排除列表也没有新增任何条目。

R1-3 [Suggestion] — 无关的不可写锁文件否决整个共享目录 (rc:4025497777)

已实现。 ci_lock_dir_writable 现在接收调用方将打开的锁文件名,只探测这些文件;目录级检查保持无条件执行,因此目录不可写时仍然总是回退。run-e2e-tests.sh 传入 docker-sandbox-daemon.lock docker-sandbox-build.lock "docker-sandbox-build-e2e-${GITHUB_SHA}.lock";清理步骤传入 docker-sandbox-daemon.lock;两处文本锚定在同一次改动中更新。新增用例:在可写的 daemon 锁旁边放一个不可写的、runner 属主的 sdk-java-tests.lock → resolver 仍输出共享目录(改动前为红,改动后转绿)。没有使用 rm 式修复——锁的身份存在于 inode 上。

R1-2 [Suggestion] — 按作业回退静默丢失协调能力;文件头夸大安全性 (rc:4025497792)

已实现「最省事的诚实修法」这一方案。 文件头现在明确说明:与同主机作业、root 的 qwen-docker-cleanup 定时器、以及 release 泳道构建互斥量之间的协调是丧失而非仅仅降级;until=24h 时限过滤保护不了被复用的镜像;清除该状态需要有人登录主机。当 GITHUB_STEP_SUMMARY 存在时,resolver 还会把同样的提示追加到其中,使降级信息出现在运行页面上,而不是只在没人打开的步骤日志里。我没有采用「主机共享回退」方案:它会引入第二个需要推理的锁命名空间,且仍然无法恢复与 root 定时器的协调(定时器以主路径为准),因此对一个本应罕见且由人工清除的状态来说,它相对诚实文档方案收益甚微。

R1-4 [Suggestion] — 签出失败时清理步骤以 127 中止 (rc:4025497806)

已实现。 由 always() 控制的清理步骤现在带有所建议的 || { ci_lock_dir="${RUNNER_TEMP:-/tmp}/qwen-code-ci-locks"; mkdir -p ...; } 保护,并注释说明这种容错只属于该步骤——测试分支继续保持大声失败。已在空工作区、bash -e 下做了行为验证:有保护时步骤 exit 0 且两条 prune 都执行;没有保护时步骤以 127 中止、prune 完全不执行。

R1-7 [Suggestion] — 大声失败分支没有测试 (rc:4025497840)

已实现。 新增用例 fails loudly when the job-private fallback cannot be created:共享目录设为 0555、RUNNER_TEMP 指向一个普通文件之下 → 断言 exit 1、stdout 为空、stderr 含 ::error::。(与建议代码有一处偏差:共享目录必须在拉起进程之前设为不可写,否则 resolver 会走主目录分支、永远到不了回退。)

R1-8 [Suggestion] — 告警断定了未经核实的 root 属主原因 (rc:4025497849)

已实现。 探测函数现在在每条返回路径上记录失败原因——cannot create <dir> / <dir> is not writable / lock file <path> is not writable——告警与错误信息都会给出出错路径。硬编码的 (unhealed root-owned leftover) 已删除。既有的锁文件回退用例现在断言告警会点名 docker-sandbox-daemon.lock;stdout 仍然恰好一行——解析出的路径——供两个调用方用 $(...) 捕获。

R1-6 [Suggestion] — exec 8> 未锚定到解析后的目录 (rc:4025497864)

已实现。 原来的裸文件名片段断言改成了完整行 exec 8>"${ci_lock_dir}/docker-sandbox-build-e2e-${GITHUB_SHA}.lock"(保留带提交标识的文件名),并新增一个行为用例:把共享目录中毒化(0555),在镜像不存在的情况下运行 docker 分支,断言 exit 0 且三个锁文件——daemon 锁、按提交协调锁、构建互斥锁——都创建在按作业私有的回退目录下。为走到构建互斥锁路径,需要 stub npm 并让 docker stub 报告镜像不存在。

变异探测

每次探测都在本轮代码树上进行,事后恢复(已按字节逐一验证恢复一致):

  • 去掉 flock stub 并用无 flock 的 PATH 农场运行 → 3 个 docker 分支用例全部变红;在同一农场下带上 stub 套件恢复全绿(即 macOS 泳道的条件)。
  • 删除 resolver 的按名探测循环 → falls back ... when a lock file is not writable 与 docker 分支锁文件用例变红——收窄没有退化成空操作。
  • 把 exec 8> 还原为 ${HOME}/.cache/qwen-code-ci 写法 → 新的 toContain 锚定与目录中毒行为用例双双变红。
  • 删除大声失败分支的 exit 1 → 新的大声失败用例变红(按该意见所述,此前全绿)。
  • 去掉按锁告警中的文件名 → 点名断言变红。
  • 删除 $GITHUB_STEP_SUMMARY 块 → 步骤摘要用例变红。
  • 删除清理步骤的 || 保护 → 保护锚定变红。

先红后绿见证(新断言对改动前脚本运行):外来锁用例、步骤摘要用例、告警点名文件断言,以及 e2e-workflow.test.js 中三处更新/新增的锚定,在改动前全部为红;脚本改动后全部转绿。

验证

  • npm run build — 通过
  • npm run typecheck — 通过
  • npm run lint — 通过(eslint;仓库的 shellcheck 泳道是 lint:all,需要下载二进制、在本沙箱中无法运行——两个 shell 脚本均通过 bash -n,并遵循仓库的 shellcheck 洁净约定)
  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/resolve-ci-lock-dir.test.js scripts/tests/e2e-workflow.test.js scripts/tests/e2e-shard-retry.test.js scripts/tests/no-ak-integration-ci.test.js — 66 通过
  • node --test .github/scripts/ci-runner-routing.test.mjs — 23 通过
  • npm run test:scripts(完整 scripts 套件)— 2499 通过,1 失败:verify-capture.test.js > preserves colour and bold independently,一个终端 ANSI 渲染断言,已通过 stash 探针在改动前的代码树上复现出完全相同的失败;它与本 PR 触碰的任何文件无关。(install-script.test.js 在 npm run build 之前也因本环境缺少 packages/audio-capture/dist 而失败,构建后已转绿。)
  • 无 flock 的 PATH 农场(/tmp/no-flock-farm:540 个条目 = 除 flock 外的全部真实 PATH 二进制)— 带 stub 时 resolver 套件全绿,即在 macOS 泳道条件下全绿。
  • npx prettier --check 检查所有触碰文件 — 全部通过。

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。

🧵 Resolved all 7 selected review thread(s). · 已关闭全部选中的 7 条评审线程。

Re-review when you have a moment. After round 10 this bot stops and leaves the PR for a human. · 有空请复审;第 10 轮后本 bot 停止并将 PR 交给人工。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.23.4

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🔀 Base updated: red check(s) [Lint & Static (ubuntu-latest, Node 22.x)] pass on current main — merged current main via update-branch; CI will re-run.

中文说明

🔀 已更新 base:红色检查 [Lint & Static (ubuntu-latest, Node 22.x)] 在当前 main 上通过 —— 已通过 update-branch 合入当前 main,CI 将重新运行。

…12006)

The prune step's 127 guard fell back to a job-private lock directory,
where flock --nonblock always "wins" against a file no other process
opens, so the labelled prune ran with its only cross-job exclusion
defeated. Probe the shared dir inline when the resolver cannot run (it
needs no checked-out file), and skip the labelled prune whenever the
daemon lock would not land on the shared dir; the dangling prune stays
unconditional because dangling images are untagged and unreferenced.

The leg now resolves the daemon lock independently of the build-family
locks, so a poisoned build lock can no longer move the daemon lock off
the shared path the prune step, the host cleanup timer and the release
lane all coordinate on.

Also pins the resolver's "cannot create" diagnostic and witnesses that
the test leg still fails loudly when no lock directory can be created.
@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🤖 Addressed the latest review feedback (round 2/10). What changed, and what I pushed back on: · 已处理最新评审反馈(第 2/10 轮)。改动内容与我反驳保留之处如下:

Autofix round summary — PR #12016 (issue #12006)

Commit: f341329b9a fix(ci): keep the docker prune exclusion on the shared daemon lock (#12006)

All seven open inline findings — two Criticals (R1-4 twice, R3-1) and two Suggestions (R1-8 twice, R2-1 twice) — are addressed in one coherent change: the daemon-lock coordination channel is now resolved and used identically by both call sites, and every new guard carries a mutation-proven witness.

R1-4 (Critical, e2e.yml:335 — job-private fallback defeats the prune exclusion) — implemented

The || { … } guard no longer invents a job-private lock directory. When the resolver cannot run (Checkout failed → 127), the step probes the shared ${HOME}/.cache/qwen-code-ci inline — it needs no checked-out file, which is why the pre-diff code worked on that path. A second guard line then blanks ci_lock_dir whenever it is not exactly the shared dir, so the labelled docker image prune --all runs only when the daemon lock genuinely lands on the shared path every leg, the host cleanup timer and the release lane coordinate on; otherwise the step takes its own else branch ("Docker cleanup skipped because the shared daemon is active"). The dangling prune moves out of the if and stays unconditional, as the finding allows — dangling images are untagged and unreferenced, so no exclusion is needed. The resolver's stderr is deliberately not silenced: ci-runner-routing.test.mjs pins doesNotMatch(run, /\/dev\/null/), so the suggested 2>/dev/null was replaced by letting mkdir's own error surface in the step log (which is what that pin's comment asks for — a failing prune must stay diagnosable). exec 9> sits inside the if condition; probed on bash: a redirection failure there becomes condition-false, not a script abort, so a dir-writable-but-lock-file-poisoned host degrades to a skip instead of a red step.

Pins moved in the same change: it('still prunes when the resolver itself cannot run') now pins the shared-dir inline probe instead of the job-private spelling; new pins assert the skip-on-non-shared guard line and the combined if [ -n … ] && exec 9>… && flock --nonblock 9 shape. New bash-driven suite prune step daemon lock discipline in scripts/tests/resolve-ci-lock-dir.test.js executes the step body extracted from the parsed e2e.yml under bash -e with stubbed docker/flock: with the resolver absent on a healthy host the lock lands on ${HOME}/.cache/qwen-code-ci/docker-sandbox-daemon.lock and the labelled prune runs; with the resolver present and the shared dir poisoned the resolver returns the job-private fallback and no labelled-prune argv is recorded (the dangling prune still runs). The lock path is witnessed through the file exec 9> creates rather than a path-recording flock stub, because a flock stub cannot see fd-target paths portably (no /proc on macOS) — the created file is the same evidence.

R3-1 (Critical, e2e.yml:333 — leg and prune resolve the daemon lock with different name sets) — implemented

The leg now resolves the coordination channel on its own:

ci_lock_dir="$(bash .github/scripts/resolve-ci-lock-dir.sh docker-sandbox-daemon.lock)"
ci_build_lock_dir="$(bash .github/scripts/resolve-ci-lock-dir.sh docker-sandbox-build.lock "docker-sandbox-build-e2e-${GITHUB_SHA}.lock")"

exec 9> (daemon read lock) stays on ${ci_lock_dir}; exec 8> (per-SHA coordinator) and exec 7> (build mutex) move to ${ci_build_lock_dir}. A poisoned build-family lock now costs only the build mutex — the leg still runs, which is this PR's goal — while the daemon lock stays on the shared path, so the prune step's flock --nonblock 9 correctly loses to a live leg and skips. The per-name probe contract (an unwritable sdk-java-tests.lock must not veto) is unchanged. Pins at e2e-workflow.test.js moved byte-for-byte in the same change, including the ordering pins (both resolutions precede every exec).

New witnesses: a resolver-level case (it.skipIf(isRoot)) poisoning only docker-sandbox-build.lock while the daemon-lock probe still resolves the shared dir, and a runDockerLeg case asserting the daemon lock file lands under the shared dir while both build-family locks land under ${RUNNER_TEMP}/qwen-code-ci-locks (image absent, so the build-mutex exec 7> path runs).

R1-8 (Suggestion, resolve-ci-lock-dir.sh:36 — cannot create message unpinned) — implemented

New case names the cause when the shared dir cannot be created: HOME points under a regular file so the primary mkdir -p fails ENOTDIR, and the case asserts stderr contains cannot create while stdout still carries the job-private fallback. ENOTDIR fails for root too, so the case needs no isRoot skip — this deviates from the suggested it.skipIf(isRoot) spelling deliberately, mirroring the same reviewer's R2-1 reasoning that ENOTDIR poisoning is uid-independent; the acceptance mutation (replace the message string) was run and reddens only this case.

R2-1 (Suggestion, run-e2e-tests.sh:41 — no witness that the leg fails loudly) — implemented

runDockerLeg now accepts home/runnerTemp env overrides. New case fails loudly when no lock directory can be created poisons both HOME and RUNNER_TEMP by ENOTDIR (regular file in the path — no isRoot skip needed) and asserts exitCode != 0, ::error:: in the output, and that the output never mentions /docker-sandbox-daemon.lock — the last assertion is what reddens if the prune step's || guard is copied into the leg "for symmetry" (the leg would otherwise open the daemon lock at the filesystem root). Because the leg now has two resolver calls, the single-line mutation is masked by the second call's failure under set -e; the realistic both-lines edit is what the probe covers, and a lane-independent text pin in e2e-workflow.test.js (not.toMatch(/resolve-ci-lock-dir\.sh[^\n]*\)"\s*\|\|/)) covers either line alone. set -euo pipefail is now also pinned (toMatch(/^set -euo pipefail$/m)). No assertion on the absence of the root-anchored FILE — on a root lane the mutant would create it, which would test the host, not the change.

Mutation probes (all run, all restored to green)

  • R1-8: problem="cannot create ${dir}" → "MUTANT-CANNOT-CREATE" → only the new case fails (1 failed / 14 passed).
  • R3-1: re-couple the leg to one three-name probe → keeps the daemon lock shared when only a build lock is poisoned fails (1/14).
  • R1-4a: restore the job-private fallback in the || { group → locks the shared dir and prunes when the resolver itself cannot run fails (1/14).
  • R1-4b: delete the skip-on-non-shared guard → skips the labelled prune when the shared dir is poisoned fails (1/14).
  • R2-1a: append || ci_lock_dir='' to both leg resolver calls → both the new leg case and the new text pin fail.
  • R2-1b: drop set -euo pipefail → both the pipefail pin and the new leg case fail.

Notes

  • The review-body item "neither discriminating input of the per-name writability probe is tested" was explicitly deferred by the reviewer under the convergence posture ("recorded, not requested in this round") and is not addressed here.
  • scripts/tests/verify-capture.test.js (preserves colour and bold independently) fails in this sandbox; reproduced identically with this round's changes stashed, so it is pre-existing and unrelated (terminal-rendering capability detection, nothing near the touched files).
  • The e2e.yml yamllint lane could not run locally (no pip/yamllint in this runner image); the change is confined to a run: |- block scalar (exempt from indentation checks via check-multi-line-strings: false), and trailing-space / comment-spacing / EOF-newline rules were verified by inspection. shellcheck (extracted from the lane's pinned tarball) reports no new findings on either touched script vs HEAD, and actionlint passes on e2e.yml.

Verification

  • npm run build — passed
  • npm run typecheck — passed
  • npm run lint — passed
  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/resolve-ci-lock-dir.test.js scripts/tests/e2e-workflow.test.js scripts/tests/e2e-shard-retry.test.js — 60 passed
  • npm run test:scripts (full scripts suite) — 2527 passed, 1 failed (verify-capture.test.js, reproduced pre-existing with this round's changes stashed)
  • node .github/scripts/ci-runner-routing.test.mjs — 23 passed, 0 failed (covers the prune-step /dev/null and filter pins)
  • actionlint (CI lane binary, on e2e.yml) — passed
  • shellcheck --check-sourced --enable=all --exclude=SC2002,SC2129,SC2310 --severity=style on both touched scripts — finding set identical to HEAD
  • bash -n .github/scripts/run-e2e-tests.sh — passed
  • npx prettier --check on the two touched test files and e2e.yml — clean after --write
  • Mutation probes — six mutations, each reddened only its intended new witness, all restored (see above)
中文说明

Autofix 本轮总结 — PR #12016(issue #12006)

提交:f341329b9a fix(ci): keep the docker prune exclusion on the shared daemon lock (#12006)

七条未结的行内意见——两条 Critical(R1-4 两次、R3-1)与两条 Suggestion(R1-8 两次、R2-1 两次)——通过一个一致的改动全部处理:daemon 锁协调通道现在由两个调用点以相同方式解析和使用,且每个新增守卫都带有经过变异验证的见证。

R1-4(Critical,e2e.yml:335——按作业私有的回退使 prune 互斥失效)——已实现

|| { … } 保护不再凭空创建按作业私有的锁目录。当 resolver 无法运行时(Checkout 失败 → 127),本步骤就地探测共享的 ${HOME}/.cache/qwen-code-ci——它不需要任何已签出的文件,这正是改动前代码在该路径上能正常工作的原因。随后第二道守卫在 ci_lock_dir 不恰好等于共享目录时将其置空,因此带标签的 docker image prune --all 只有在 daemon 锁真正落在共享路径上时才会执行——该路径是所有测试分支、主机清理定时器与 release 泳道共同协调的位置;否则本步骤进入自带的 else 分支("Docker cleanup skipped because the shared daemon is active")。按该意见允许的方式,悬空镜像 prune 移出 if 并保持无条件执行——悬空镜像无标签、无引用,不需要互斥。resolver 的 stderr 刻意不加静音:ci-runner-routing.test.mjs 锚定了 doesNotMatch(run, /\/dev\/null/),因此建议中的 2>/dev/null 被替换为让 mkdir 自身的错误直接显示在步骤日志中(这正是该锚定注释所要求的——失败的 prune 必须保持可诊断)。exec 9> 位于 if 条件内;已在 bash 上实测:该处的重定向失败会变成条件为假,而不是中止脚本,因此「目录可写但锁文件被污染」的主机会优雅降级为跳过,而不是红步骤。

同一次改动中移动的锚定:it('still prunes when the resolver itself cannot run') 现在锚定共享目录的就地探测,而不再是按作业私有的写法;新增锚定断言了「非共享即跳过」守卫行以及 if [ -n … ] && exec 9>… && flock --nonblock 9 的组合形态。scripts/tests/resolve-ci-lock-dir.test.js 新增 bash 驱动套件 prune step daemon lock discipline:从解析后的 e2e.yml 中取出步骤正文,在 bash -e 下用 stub 的 docker/flock 执行:resolver 缺失且主机健康时,锁落在 ${HOME}/.cache/qwen-code-ci/docker-sandbox-daemon.lock 且带标签 prune 执行;resolver 存在且共享目录中毒时,resolver 返回按作业私有的回退,且不会记录任何带标签 prune 的参数(悬空 prune 仍会执行)。锁路径通过 exec 9> 所创建的文件来见证,而不是用记录路径的 flock stub——因为 flock stub 无法可移植地看到文件描述符的目标路径(macOS 没有 /proc),创建出的文件即是同等证据。

R3-1(Critical,e2e.yml:333——分支与 prune 用不同的锁名集合解析 daemon 锁)——已实现

分支现在独立解析协调通道:

ci_lock_dir="$(bash .github/scripts/resolve-ci-lock-dir.sh docker-sandbox-daemon.lock)"
ci_build_lock_dir="$(bash .github/scripts/resolve-ci-lock-dir.sh docker-sandbox-build.lock "docker-sandbox-build-e2e-${GITHUB_SHA}.lock")"

exec 9>(daemon 读锁)保持使用 ${ci_lock_dir};exec 8>(按 SHA 协调锁)与 exec 7>(构建互斥锁)改用 ${ci_build_lock_dir}。中毒的构建族锁现在只损失构建互斥量——分支仍能运行,这正是本 PR 的目标——而 daemon 锁保持在共享路径上,因此 prune 步骤的 flock --nonblock 9 会正确地输给存活的分支并跳过。按名探测的契约(不可写的 sdk-java-tests.lock 不得否决)保持不变。e2e-workflow.test.js 中的锚定在同一次改动中逐字节移动,包括顺序锚定(两次解析都先于所有 exec)。

新增见证:一个 resolver 级用例(it.skipIf(isRoot))只毒化 docker-sandbox-build.lock,而 daemon 锁探测仍解析到共享目录;以及一个 runDockerLeg 用例,断言 daemon 锁文件落在共享目录下、两把构建族锁落在 ${RUNNER_TEMP}/qwen-code-ci-locks 下(镜像缺失,因此会走到构建互斥锁 exec 7> 的路径)。

R1-8(Suggestion,resolve-ci-lock-dir.sh:36——cannot create 消息未被锚定)——已实现

新增用例 names the cause when the shared dir cannot be created:HOME 指向一个普通文件之下,使主目录的 mkdir -p 以 ENOTDIR 失败,用例断言 stderr 含有 cannot create,同时 stdout 仍携带按作业私有的回退路径。ENOTDIR 对 root 同样失败,因此该用例不需要 isRoot 跳过——这里有意偏离了建议的 it.skipIf(isRoot) 写法,沿用同一评审者在 R2-1 中给出的「ENOTDIR 毒化与 uid 无关」的推理;验收变异(替换该消息字符串)已运行,且只有这个新用例变红。

R2-1(Suggestion,run-e2e-tests.sh:41——缺少「分支大声失败」的见证)——已实现

runDockerLeg 现在接受 home/runnerTemp 环境变量覆盖。新增用例 fails loudly when no lock directory can be created 以 ENOTDIR 方式同时毒化 HOME 与 RUNNER_TEMP(路径中放一个普通文件——无需 isRoot 跳过),断言 exitCode != 0、输出含 ::error::、且输出从不提及 /docker-sandbox-daemon.lock——最后一条断言正是当有人「为了对称」把 prune 步骤的 || 保护复制进分支时会让用例变红的东西(否则分支会去文件系统根目录打开 daemon 锁)。由于分支现在有两条 resolver 调用,单行变异在 set -e 下会被第二条的失败掩盖;变异探针覆盖的是现实中两行同改的情形,另有 e2e-workflow.test.js 中与泳道无关的文本锚定(not.toMatch(/resolve-ci-lock-dir\.sh[^\n]*\)"\s*\|\|/))覆盖任意单行。set -euo pipefail 现在也被锚定(toMatch(/^set -euo pipefail$/m))。没有对根路径文件「不存在」的断言——在 root 泳道上变异体会创建它,那测的就是主机而不是改动。

变异探针(全部运行,且均已恢复为绿)

  • R1-8:problem="cannot create ${dir}" → "MUTANT-CANNOT-CREATE" → 只有新用例失败(1 失败 / 14 通过)。
  • R3-1:把分支重新耦合为一次三名探测 → keeps the daemon lock shared when only a build lock is poisoned 失败(1/14)。
  • R1-4a:在 || { 组中恢复按作业私有回退 → locks the shared dir and prunes when the resolver itself cannot run 失败(1/14)。
  • R1-4b:删除「非共享即跳过」守卫 → skips the labelled prune when the shared dir is poisoned 失败(1/14)。
  • R2-1a:给分支的两条 resolver 调用都追加 || ci_lock_dir='' → 新分支用例与新文本锚定同时失败。
  • R2-1b:删除 set -euo pipefail → pipefail 锚定与新分支用例同时失败。

备注

  • 评审正文中的「per-name 可写性探测的两个区分性输入都未被测试」一条,已由评审者在收敛姿态下明确延后(「已记录,本轮不要求修改」),本轮不处理。
  • scripts/tests/verify-capture.test.js(preserves colour and bold independently)在本沙箱中失败;在把本轮改动 stash 后仍以同样方式复现,因此是既有失败且与本改动无关(终端渲染能力探测,与所涉文件无关)。
  • e2e.yml 的 yamllint 泳道无法在本地运行(本 runner 镜像没有 pip/yamllint);改动完全位于 run: |- 块标量内(经 check-multi-line-strings: false 豁免缩进检查),行尾空格、注释空格、文件末尾换行等规则已通过检查确认。shellcheck(从该泳道固定的 tarball 解出)在两个被改脚本上的报告与 HEAD 完全一致,actionlint 对 e2e.yml 通过。

Verification

  • npm run build — 通过
  • npm run typecheck — 通过
  • npm run lint — 通过
  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/resolve-ci-lock-dir.test.js scripts/tests/e2e-workflow.test.js scripts/tests/e2e-shard-retry.test.js — 60 通过
  • npm run test:scripts(完整 scripts 套件)— 2527 通过,1 失败(verify-capture.test.js,已在 stash 本轮改动后复现为既有失败)
  • node .github/scripts/ci-runner-routing.test.mjs — 23 通过,0 失败(覆盖 prune 步骤的 /dev/null 与 filter 锚定)
  • actionlint(CI 泳道二进制,对 e2e.yml)— 通过
  • shellcheck --check-sourced --enable=all --exclude=SC2002,SC2129,SC2310 --severity=style 对两个被改脚本 — 发现集合与 HEAD 完全一致
  • bash -n .github/scripts/run-e2e-tests.sh — 通过
  • npx prettier --check 对两个被改测试文件与 e2e.yml — --write 后干净
  • 变异探针 — 六个变异,每个只让其对应的新见证变红,且均已恢复(见上文)

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。

🧵 Resolved all 4 selected review thread(s). · 已关闭全部选中的 4 条评审线程。

Re-review when you have a moment. After round 10 this bot stops and leaves the PR for a human. · 有空请复审;第 10 轮后本 bot 停止并将 PR 交给人工。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.24.0

)

Address review feedback on the lock-dir fallback:

- The prune step blanked a job-private resolver result but still logged
  "skipped because the shared daemon is active", and its resolver-absent
  inline probe checked only the directory, so a root-owned 0400 lock file
  failed `exec 9>` with EACCES inside the if-condition (no bash -e abort)
  and the step exited 0 with that false line as its only output. The step
  now probes the lock file inline, carries the cause it actually observed
  (naming the lock dir or file), and announces the unconditional dangling
  prune separately so "skipped" scopes to the labelled one. Its resolver
  call blanks GITHUB_STEP_SUMMARY: the step discards a job-private result
  and opens no lock in it, so the fallback banner would lie there (the
  leg's calls keep it).
- The resolver's fallback warning and summary claimed "docker build/prune
  coordination is lost" even when only a build-family lock fell back;
  both now scope the loss to the locks the call actually probed ($*).
- Tests: pin the unwritable-dir cause text at both reach points (and fix
  the comment that claimed it was pinned), add the success-direction
  witness driving the resolver and the step's shared-dir guard together
  so a one-token drift between them reddens, and assert the opened-locks
  == probed-locks invariant as a derived set so a lock added to an exec
  line without a probe-list entry fails the suite.
@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🤖 Addressed the latest review feedback (round 3/10). What changed, and what I pushed back on: · 已处理最新评审反馈(第 3/10 轮)。改动内容与我反驳保留之处如下:

Addressed review feedback — PR #12016 round

All five inline findings from the latest automated review were verified against the code, reproduced where they claimed wrong behavior, and resolved in commit d75b0bb00b (fix(ci): scope the docker cleanup diagnostics to what was probed (#12006)). --conflict false: no merge of main was needed or performed. No test was deleted or weakened — the changed files' assertion counts only grew, so no test-weakening.json is required.

Dispositions

[rc:4031664803] R4-1 — prune step asserted causes it never verified — RESOLVED

Verified before implementing: on the resolver-absent path the inline probe tested only [ -w ] on the directory, so a root-owned 0400 daemon lock failed exec 9> with EACCES inside the if condition (no bash -e abort) and the step exited 0 printing only the false Docker cleanup skipped because the shared daemon is active; on the resolver-fallback path the same message claimed an active daemon while the step's own guard had just discarded the resolver's job-private result; and the resolver's summary banner claimed this job locks in the job-private <dir> for the one caller that opens no lock there.

Changes in .github/workflows/e2e.yml:

  • The resolver-absent inline fallback now probes the daemon lock file ([ -e ] && [ ! -w ]), not just the directory, and records skip_reason naming the offending path (.../docker-sandbox-daemon.lock is not writable by this job / cannot create a writable ...).
  • The shared-dir guard only fires when the resolver returned a non-shared dir and carries its own reason; the skip line is now Docker cleanup skipped: <observed reason>; clearing ${HOME}/.cache/qwen-code-ci needs a human on this host. The shared daemon is active message survives only on the branch where flock --nonblock actually lost.
  • The step's resolver call is GITHUB_STEP_SUMMARY= bash ... so the fallback banner is suppressed for exactly this discarding caller; the leg's two calls keep the banner.
  • Running the unconditional dangling prune is announced separately, so a "skipped" line scopes to the labelled prune.

Tests: retargeted skips the labelled prune when the shared dir is poisoned (skip line must name the lock dir and needs a human; not.toContain('shared daemon is active')), added names the unwritable daemon lock when the resolver cannot run (withResolver: false + 0o400 lock in a writable shared dir) and writes no fallback banner for the caller that discards the result (GITHUB_STEP_SUMMARY set; summary stays empty while the labelled prune is skipped). The e2e-workflow.test.js text pins moved with the restructure (blanked resolver call, != guard, elif exec 9>).

[rc:4031664817] R4-2 — no success-direction witness for the resolver/step path agreement — RESOLVED

Verified: runPruneStep(world, { withResolver: true }) was used exactly once and asserted the prune does not run; the healthy path had no behavioral witness. Added locks the shared dir and prunes when the resolver is healthy: clean world, resolver present, asserting exitCode === 0, the daemon lock file exists under the shared ${HOME}/.cache/qwen-code-ci, and dockerArgv contains image prune --all. Kept the string-identity guard (the finding's primary ask is the witness; the flag/exit-status redesign was explicitly the alternative). Guard deletion is caught by the existing poisoned-dir case, so both mutations now have behavioral witnesses.

[rc:4031664822] R4-3 — fallback messages claimed prune coordination loss on build-only fallback — RESOLVED

Verified: the shared template printed docker build/prune coordination on this host is lost and named the qwen-docker-cleanup timer even when only a build-family lock fell back — the state this PR's split made reachable. .github/scripts/resolve-ci-lock-dir.sh now scopes both the ::warning:: and the step-summary sentence to the locks the call actually probed (for: $*; cross-job coordination on these locks is lost for this job), keeping ${problem}, the lock-dir fallback heading, both paths, and the does-not-heal sentence. Extended the leg-level case keeps the daemon lock shared when only a build lock is poisoned with output/summary assertions: names docker-sandbox-build.lock, never matches /prune/. Left the daemon-only resolver case untouched (its stderr is asserted empty there, so a not.toMatch would be vacuous).

[rc:4031664833] R4-4 — unwritable-dir cause text unpinned — RESOLVED

Verified: cause 2 (${dir} is not writable) had no message-text pin and the comment on the cannot create case wrongly claimed it was "pinned above". Added expect(result.stderr).toContain('is not writable') to falls back when the shared dir itself is not writable and to the ::error:: case, and reworded the comment to point at the unwritable-lock pin above and the unwritable-dir pin below.

[rc:4031664844] R4-5 — opened-locks == probed-locks invariant pinned only as literals — RESOLVED

Verified: a lock added to an exec N> line without extending the resolver argument list left every gate green. Added probes exactly the lock files each call site opens in e2e-workflow.test.js: it extracts the opened set from exec \d+>"${ci_lock_dir}/…" / exec \d+>"${ci_build_lock_dir}/…" and the probed set from the matching resolve-ci-lock-dir.sh … invocation (per variable), comparing sorted string sets with ${GITHUB_SHA} unexpanded on both sides — for run-e2e-tests.sh (both variables) and the prune step in e2e.yml.

Mutation probes (guard removed → matching case reddens → restored → green)

Probe Result
Dropped needs a human clause from the skip echo skips the labelled prune when the shared dir is poisoned red
Reverted inline probe to dir-only names the unwritable daemon lock when the resolver cannot run red
Dropped GITHUB_STEP_SUMMARY= prefix writes no fallback banner … red (+ the two moved text pins)
Renamed resolver primary literal locks the shared dir and prunes when the resolver is healthy red (among others)
Deleted the shared-dir guard poisoned-dir and banner cases red
Reverted the R4-3 warning text keeps the daemon lock shared when only a build lock is poisoned red
Mutated cause-2 text to cannot create both is not writable pins red
Added unprobed exec 6> lock to run-e2e-tests.sh / to the prune step probes exactly the lock files each call site opens red

Notes

  • The review's two convergence-deferred items (2>/dev/null on the primary mkdir -p; the header's per-job RUNNER_TEMP claim) were explicitly "recorded, not requested in this round" and are untouched.
  • scripts/tests/ci-runner-routing.test.mjs, cited in R4-1's constraint list, does not exist in this repository; the substance of the constraint (both prune commands, until=24h, || echo "::warning::, no /dev/null silencing) is preserved — the resolver is muted per-call via the GITHUB_STEP_SUMMARY= env prefix, not a redirect.
  • Two unrelated scripts-suite failures reproduce identically with this round's changes stashed on the pre-round HEAD: verify-capture helper > preserves colour and bold independently (ANSI rendering) and standalone release packaging > does not package audio-capture test artifacts (create-standalone-package.js smoke run). Neither file reads the scripts this PR touches; both are pre-existing on this runner.

Verification

  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/resolve-ci-lock-dir.test.js scripts/tests/e2e-workflow.test.js — 56 passed (was 52 pre-round; +4 new cases)
  • npx vitest run … scripts/tests/e2e-shard-retry.test.js (the third reader of run-e2e-tests.sh) — 8 passed
  • npm run test:scripts (full scripts suite) — 2528 passed, 2 failed; both failures reproduced identically on the pre-round tree with this round stashed (pre-existing, unrelated — see Notes)
  • npm run build — passed
  • npm run typecheck — passed
  • npm run lint — passed
  • npx prettier --check on the two changed test files and e2e.yml — passed (.sh has no prettier parser)
  • bash -n on resolve-ci-lock-dir.sh, run-e2e-tests.sh, and the extracted prune-step body — passed
  • Mutation probes — 8/8 reddened the intended witness and were restored to green (table above)
中文说明

已处理评审反馈 —— PR #12016 本轮

最新一轮自动评审的 5 条行内发现全部先对照代码核实、对声称行为错误的部分做了复现,并在提交 d75b0bb00b(fix(ci): scope the docker cleanup diagnostics to what was probed (#12006))中解决。--conflict false:不需要也未执行与 main 的合并。没有删除或削弱任何测试——被改动文件的断言数量只增不减,因此无需 test-weakening.json。

逐条处置

[rc:4031664803] R4-1 —— 清理步骤断言了从未验证的成因 —— 已解决

实现前已核实:在解析器缺失路径上,就地探测只对目录做 [ -w ],于是 root 属主的 0400 daemon 锁会让 exec 9> 在 if 条件内以 EACCES 失败(bash -e 不会中止),步骤以 0 退出且唯一输出就是那句错误的 Docker cleanup skipped because the shared daemon is active;在解析器回退路径上,同一句在步骤刚丢弃解析器的 job-private 结果后仍声称 daemon 活跃;而解析器的摘要横幅会为这个根本不在该目录加锁的调用点声称「本作业锁进了 job-private <dir>」。

.github/workflows/e2e.yml 的改动:

  • 解析器缺失时的就地回退现在探测 daemon 锁文件([ -e ] && [ ! -w ]),而不只是目录,并用 skip_reason 记录肇事路径(.../docker-sandbox-daemon.lock is not writable by this job / cannot create a writable ...)。
  • 共享目录守卫只在解析器返回了非共享目录时触发并携带自己的原因;跳过行变为 Docker cleanup skipped: <实际观察到的原因>; clearing ${HOME}/.cache/qwen-code-ci needs a human on this host。shared daemon is active 只保留在 flock --nonblock 真正抢锁失败的分支。
  • 本步骤的解析器调用改为 GITHUB_STEP_SUMMARY= bash ...,恰好为这个丢弃结果的调用点抑制回退横幅;测试分支的两处调用保留横幅。
  • 单独打印 Running the unconditional dangling prune,使「skipped」只指带标签的清理。

测试:改写了 skips the labelled prune when the shared dir is poisoned(跳过行必须点名锁目录并含 needs a human;not.toContain('shared daemon is active')),新增 names the unwritable daemon lock when the resolver cannot run(withResolver: false + 可写共享目录中的 0o400 锁)与 writes no fallback banner for the caller that discards the result(设置 GITHUB_STEP_SUMMARY;摘要保持为空且带标签清理被跳过)。e2e-workflow.test.js 的文本锚定随重构一并移动(带置空前缀的解析器调用、!= 守卫、elif exec 9>)。

[rc:4031664817] R4-2 —— 解析器与步骤路径约定缺少成功方向见证 —— 已解决

已核实:runPruneStep(world, { withResolver: true }) 只被用过一次且断言清理不执行;健康路径没有行为见证。新增 locks the shared dir and prunes when the resolver is healthy:干净环境、解析器在场,断言 exitCode === 0、daemon 锁文件落在共享的 ${HOME}/.cache/qwen-code-ci 下、且 dockerArgv 含 image prune --all。保留字符串相等守卫(该发现的首要诉求就是补见证;改旗标/退出状态的方案本就是可选的另一条路)。删除守卫会被既有的目录中毒用例捕获,因此两个变异方向现在都有行为见证。

[rc:4031664822] R4-3 —— 仅构建族回退也声称清理协调丢失 —— 已解决

已核实:共享模板在只有构建族锁回退时也打印 docker build/prune coordination on this host is lost 并点名 qwen-docker-cleanup 定时器——这正是本 PR 的拆分新造出的状态。.github/scripts/resolve-ci-lock-dir.sh 现在把 ::warning:: 与摘要句都限定到本次调用实际探测的锁(for: $*; cross-job coordination on these locks is lost for this job),并保留 ${problem}、lock-dir fallback 标题、两条路径以及「不会自愈」那句。扩展分支级用例 keeps the daemon lock shared when only a build lock is poisoned:输出与摘要都必须点名 docker-sandbox-build.lock、且不匹配 /prune/。按发现要求未改动 daemon-only 的解析器用例(那里已断言 stderr 为空,再加 not.toMatch 是空断言)。

[rc:4031664833] R4-4 —— 目录不可写成因的文本未被锚定 —— 已解决

已核实:成因 2(${dir} is not writable)没有任何消息文本锚定,而 cannot create 用例的注释错误地声称它「已在上方锚定」。在 falls back when the shared dir itself is not writable 与 ::error:: 用例各加 expect(result.stderr).toContain('is not writable'),并把注释改写为指向上方的不可写锁锚定与下方的不可写目录锚定。

[rc:4031664844] R4-5 —— 「打开的锁 == 探测的锁」不变量只有字面量锚定 —— 已解决

已核实:给 exec N> 行加锁而不扩展解析器参数名单时所有门禁全绿。在 e2e-workflow.test.js 新增 probes exactly the lock files each call site opens:从 exec \d+>"${ci_lock_dir}/…" / exec \d+>"${ci_build_lock_dir}/…" 提取打开集合,从同一变量对应的 resolve-ci-lock-dir.sh … 调用提取探测集合,两侧保留 ${GITHUB_SHA} 不展开、按排序后的字符串集合比较——覆盖 run-e2e-tests.sh(两个变量)与 e2e.yml 的清理步骤。

变异探针(拆除守卫 → 对应用例变红 → 还原 → 复绿)

探针 结果
删除跳过行中的 needs a human 从句 skips the labelled prune when the shared dir is poisoned 变红
就地探测退回只查目录 names the unwritable daemon lock when the resolver cannot run 变红
删除 GITHUB_STEP_SUMMARY= 前缀 writes no fallback banner … 变红(外加两条随动的文本锚定)
改名解析器的 primary 字面量 locks the shared dir and prunes when the resolver is healthy 变红(及其他)
删除共享目录守卫 目录中毒用例与横幅用例变红
还原 R4-3 的警告文本 keeps the daemon lock shared when only a build lock is poisoned 变红
把成因 2 文本改为 cannot create 两处 is not writable 锚定变红
给 run-e2e-tests.sh / 清理步骤加未探测的 exec 6> 锁 probes exactly the lock files each call site opens 变红

说明

  • 评审中两条按收敛姿态延后的条目(主 mkdir -p 上的 2>/dev/null;文件头关于 RUNNER_TEMP 按作业隔离的说法)明确写着「已记录,本轮不要求修改」,未动。
  • R4-1 约束清单提到的 scripts/tests/ci-runner-routing.test.mjs 在本仓库不存在;该约束的实质(两条清理命令、until=24h、|| echo "::warning::、禁止用 /dev/null 静音解析器)均已保留——解析器是按调用点用 GITHUB_STEP_SUMMARY= 环境前缀静默的,不是重定向。
  • 两个与本改动无关的 scripts 套件失败,在本轮改动 stash 后的轮前 HEAD 上逐字复现:verify-capture helper > preserves colour and bold independently(ANSI 渲染)与 standalone release packaging > does not package audio-capture test artifacts(create-standalone-package.js 冒烟运行)。两者都不读本 PR 触及的脚本,属于本运行器上的既有失败。

验证

  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/resolve-ci-lock-dir.test.js scripts/tests/e2e-workflow.test.js —— 56 通过(轮前 52;新增 4 例)
  • npx vitest run … scripts/tests/e2e-shard-retry.test.js(run-e2e-tests.sh 的第三个读取方)—— 8 通过
  • npm run test:scripts(完整 scripts 套件)—— 2528 通过,2 失败;两个失败在 stash 本轮改动后的轮前树上逐字复现(既有且无关,见「说明」)
  • npm run build —— 通过
  • npm run typecheck —— 通过
  • npm run lint —— 通过
  • 对两个改动的测试文件与 e2e.yml 执行 npx prettier --check —— 通过(.sh 无 prettier 解析器)
  • 对 resolve-ci-lock-dir.sh、run-e2e-tests.sh 及提取出的清理步骤体执行 bash -n —— 通过
  • 变异探针 —— 8/8 均使预期见证变红并已还原复绿(见上表)

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。

🧵 Resolved all 5 selected review thread(s). · 已关闭全部选中的 5 条评审线程。

Re-review when you have a moment. After round 10 this bot stops and leaves the PR for a human. · 有空请复审;第 10 轮后本 bot 停止并将 PR 交给人工。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.24.0

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

⏸️ AutoFix paused by a review convergence signal: 3 consecutive automated review round(s) signalled machine-readable non-convergence (codes observed since the last maintainer response, or the window start if none: batch-fixes, root-cause-triage, stem-surface) — the reviewer keeps finding new defects at a rate that is not falling while the loop keeps widening the diff, so another automatic round is unlikely to converge this PR. The loop resumes once a maintainer responds on this PR (a review or comment counts, and steers the next rounds), and pauses again if the signal persists for 3 more round(s). Alternatives: split the recurring cluster into its own PR, batch the remaining fixes into one push, comment @qwen-code /retry to re-arm a fresh window, or @qwen-code /takeover to take it over with the raised cap. While paused, new feedback and base conflicts stay unhandled.

中文说明

⏸️ AutoFix 已因评审收敛信号暂停:3 轮自动评审连续发出机器可读的不收敛信号(自上次维护者响应以来观察到的信号码;若无响应则自窗口开始:batch-fixes, root-cause-triage, stem-surface)——评审仍在以不降的速率发现新缺陷,而循环在继续扩大 diff,再跑一轮自动修复难以收敛本 PR。维护者在本 PR 上作出回应后循环自动恢复(评论或评审均可,并将作为后续轮次的指引);若信号再持续 3 轮会再次暂停。可选做法:把反复出问题的簇拆成独立 PR、把剩余修复攒成一批一次推送、评论 @qwen-code /retry 重开计数窗口、或评论 @qwen-code /takeover 以更高轮次上限接管。暂停期间,新反馈与 base 冲突不会被处理。

@wenshao

wenshao commented Sep 20, 2026

Copy link
Copy Markdown
Collaborator

@qwen-code /triage

@wenshao
wenshao enabled auto-merge September 20, 2026 03:05
@wenshao

wenshao commented Sep 20, 2026

Copy link
Copy Markdown
Collaborator

Maintainer verification — built the real environment locally and drove this change through it

Verified at head 1cf3bc952b45d9857780b806674ee7564a18f8da (net effect against today's origin/main 8589f713: 6 files, +958/−11, matching the PR description). Verdict: this is safe to land, with one residual risk the standing Critical describes correctly, and two sentences in the PR description that need correcting.

How it was verified

Not by re-reading the diff. A real Linux host was reconstructed and the three step bodies were executed on it, byte-for-byte as the workflow ships them:

  • Container: Debian bookworm, real flock (util-linux 2.38.1), uid 2001 github-runner, no passwordless sudo — the hk1/hk2 shape, confirmed in-rig (sudo: a password is required).
  • Step bodies: Restore workspace ownership, Run E2E tests and Prune dangling docker images extracted from e2e.yml by YAML parse (no rewriting), run in order under GitHub's default bash --noprofile --norc -e -o pipefail (e2e.yml sets no shell:/defaults: override).
  • Arms: origin/main vs the PR merged into it, same rig, same worlds.
  • Not real: docker/npm/npx are argv-recording shims, so "what the job asked docker to do" is exact while the image build and vitest payload are not. The one claim that needs a real daemon — R7-1's image loss — was re-run against real Docker 29.5.2 separately.

1. The failure reproduces on main, and the PR fixes it

repro

The main arm emits the production log of run 35069321648 line for line, including the line number: run-e2e-tests.sh: line 46: …/docker-sandbox-daemon.lock: Permission denied, then the prune step's own line 2: copy, both exit 1, no test executed. On the PR arm the same host state gives leg rc=0 · prune rc=0.

matrix

Two details of the design hold up under test, and neither is pinned by reading alone:

  • A poisoned build lock moves only the build family — the daemon lock stays shared, so the prune exclusion survives (main dies at line 65 in the same world).
  • A poisoned sdk-java-tests.lock vetoes nothing, because each call probes only its own names.

2. The incident is wider than the linked issue

Pulling every sandbox:docker leg from the e2e workflow between 2026-09-16 and 2026-09-20 and checking each failure's log: four legs died of this exact EACCES signature on 2026-09-16, across two hosts — hk1-24 00:27Z, hk2-29 04:26Z, hk2-21 04:39Z, hk2-2 07:45Z (the one #12006 tracks). Every other failure in the window was a different cause.

From 2026-09-17 04:11Z onward the same hosts run the leg green, so the poisoned state ended within about 21 hours. Whether a human cleared it or something else did is not visible from the logs — worth noting because the resolver header asserts "The state does not heal itself: clearing it needs a human on the host," and production neither confirms nor refutes that.

3. On a healthy host the change is not quite byte-for-byte

With no sibling leg running, every recorded stream is identical between the arms except one added stdout line — same docker argv, same locks, same exit codes. But with a sibling leg live, the arms diverge:

main PR
labelled prune skipped skipped
dangling prune skipped runs

The PR moves docker image prune --force --filter until=24h out from under the daemon lock, so it now runs while a sibling leg is mid-test. This is defensible — dangling images are untagged and unreferenced, the age gate is 24h, and the host's own root timer already prunes them without the lock (qwen-docker-cleanup.sh:44) — but the description's "a healthy host is byte-for-byte the old behavior" is not exact, and the same sentence appears in the Chinese section ("健康主机与旧行为逐字节一致").

4. The standing Critical (R7-1) — confirmed, and it is not merely theoretical

timer

Driving the real qwen-docker-cleanup.sh as root against a live leg:

  • Healthy host: the leg holds the shared daemon lock → timer prints skipped: sandbox daemon lock busy, labelled prune never runs.
  • Poisoned host, PR arm: the leg is in $RUNNER_TEMP, the timer still resolves the shared path, is not skipped by the poison file (:53 only skips when the lock is absent), takes it uncontended as root, and runs the labelled prune against a daemon with a live leg's docker work in flight.

Against real Docker: an image the leg has just accepted via docker image inspect survives the timer's real command at until=24h (negative control, seconds old) and is deleted by the identical command once the threshold puts it past the gate (positive control — threshold shifted rather than waiting 24h), after which docker run fails No such image. R7-1's chain is real end to end.

Where R7-1's premise is only half-supported. It rests on the poison being a root-owned lock file inside a runner-owned directory, which the runner can rm -f with no sudo. Both shapes were built:

host state runner's rm -f resolver after
runner-owned dir + root-owned 0400 lock rc=0, poison gone back on the shared dir
root-owned dir + root-owned 0400 lock rc=1, Permission denied still falls back

And the two shapes are indistinguishable in production: the heal warning, the leg's line 46 error and the prune step's line 2 error are byte-identical between them. So #12006's log cannot tell us which shape hk2-2 had, and R7-1's repair covers one of the two while this PR's fallback covers both.

R7-1's proposed repair was implemented and measured (narrow rm -f of only the probed-unwritable names, re-test, fall back on failure):

world leg daemon lock poison afterwards
root-owned lock file 0 stays SHARED cleared, runner-owned 0644
root-owned directory 0 job-private (correct) untouched
healthy 0 SHARED unchanged

It works, it does not weaken the unrepairable case, and it costs 3 existing cases that currently assert "unwritable lock file → fallback" and would need re-pointing at the unrepairable shape. That is a real follow-up, not a reason to hold this PR: on the shapes it cannot repair, the fallback is still the only thing standing between the host and a dead leg.

5. The other open findings

mutations

  • R5-1 (open) — confirmed. On a poisoned host the prune step emits ::warning::… using job-private lock dir … on stderr while its stdout says Docker cleanup skipped and it opens no lock at all. Two signals, contradicting each other, in the same step.
  • R5-2 (open) — confirmed by mutation. Deleting ci_lock_dir='' from the inline fallback's else passes all 56 tests, and on a real host flips the message from cannot create a writable <dir>; … needs a human to Docker cleanup skipped because the shared daemon is active — a cause no probe verified. The branch itself behaves correctly as shipped (verified in all three resolver-absent worlds); only the coverage is missing.
  • R9-1 — confirmed. When even the fallback dir cannot be created, the resolver's ::error:: still reaches the annotation list while the step recovers to rc=0.
  • R9-2 — confirmed, precisely. Deleting the unconditional dangling prune is killed by 2 cases, both on the poisoned path; the healthy-path cases do not pin it.
  • R9-3 — confirmed. The two expect(output).not.toMatch(/prune/) assertions cannot fail: none of the resolver's six output templates contains the word.
  • R9-4 — confirmed. The permanently-disabled labelled prune reaches no annotation and no run-page summary; it exists only as one stdout line.

Mutation coverage of what the PR actually protects is otherwise strong — 7 of 8 mutations killed, including every resolver-core one (3–9 cases each).

6. Review gaps this closes

The last review disclosed four gaps. Three are now closed:

  • macOS lane: resolve-ci-lock-dir.test.js + e2e-workflow.test.js = 56/56 on Darwin (bash 5.3.15), and 18/18 under Apple's own bash 3.2.57 (/bin/bash, not a GNU 3.2 container).
  • shellcheck (repo's exact invocation, pinned 0.11.0): the new script produces one SC2292 style note, the same class the existing run-e2e-tests.sh already carries ~30 of. Nothing new.
  • yamllint (pinned 1.35.1, --format github): clean on both arms. The PR lists this as unvalidated.
  • actionlint (pinned 1.7.12): clean on both arms with the repo's flags (-shellcheck= disables the embedded-shell pass). Note for accuracy: with default flags the PR arm reports SC1007 on the intentional GITHUB_STEP_SUMMARY= bash … prefix assignment while main is silent — a false positive that the repo's own gate never sees, but "actionlint parses the file cleanly" is true only under the repo's flags.

Recommendation

Merge. The change fixes a failure that is reproducible, recently recurrent across two pool hosts, and fatal to the leg before a single test runs; on a healthy host it costs one log line and one prune that the host's own timer already performs unlocked. The standing Critical is accurate about the cost on a poisoned host, but the alternative on such a host is a leg that cannot run at all.

Two follow-ups worth filing rather than blocking on:

  1. R7-1's narrow repair (measured above): it converts a permanent degradation into self-healing for the shape Main CI failed: E2E Tests on aff26a8f36df #12006's triage describes, at the price of re-pointing 3 cases.
  2. R5-1 / R9-4 diagnostics: on a poisoned host the operator currently gets a warning naming a lock the step never opened, and no annotation at all for the labelled prune being permanently off.

Two corrections to make in the description before merge: the "byte-for-byte" sentence (§3, both languages), and Fixes #12006 — that issue was closed on 2026-09-18 as superseded by #11991, which is where this belongs.

中文说明

维护者验证 —— 在本地搭建真实环境并把本改动放进去跑

验证对象为 head 1cf3bc952b(相对今天的 origin/main 8589f713 的净效果:6 个文件、+958/−11,与 PR 描述一致)。结论:可以合并;未决 Critical 对代价的描述是准确的;PR 描述中有两处需要更正。

验证方式

不是靠读 diff,而是重建了一台真实的 Linux 主机,并按工作流原样逐字执行三个步骤体:

  • 容器:Debian bookworm,真 flock(util-linux 2.38.1),uid 2001 github-runner,无免密 sudo——即 hk1/hk2 的形态(装置内实测 sudo: a password is required)。
  • 步骤体:用 YAML parse 从 e2e.yml 逐字取出 Restore workspace ownership、Run E2E tests、Prune dangling docker images(不改写),按顺序在 GitHub 的默认 shell bash --noprofile --norc -e -o pipefail 下运行(e2e.yml 没有任何 shell:/defaults: 覆盖)。
  • 两臂:origin/main 与「PR 合入 main 后」的结果,同一装置、同样的世界。
  • 不真实的部分:docker/npm/npx 是记录 argv 的垫片,因此「作业请求 docker 做了什么」是精确的,而镜像构建与 vitest 负载不是。唯一需要真守护进程的论断(R7-1 的镜像丢失)另外用真实 Docker 29.5.2 复跑。

1. 故障在 main 上复现,PR 修好了它

main 臂逐行重现了 run 35069321648 的生产日志,连行号都一致:run-e2e-tests.sh: line 46: …/docker-sandbox-daemon.lock: Permission denied,随后是 prune 步骤自己的 line 2:,两者都 exit 1,一个测试都没跑。同样的主机状态在 PR 臂上是 leg rc=0 · prune rc=0。

设计中的两个细节经测试成立,且都不是读代码能钉住的:

  • 被污染的 build 锁只挪动 build 家族——daemon 锁仍在共享目录,prune 互斥得以保留(同样世界里 main 死在 line 65)。
  • 被污染的 sdk-java-tests.lock 否决不了任何东西,因为每次调用只探测自己的锁名。

2. 事故范围比关联 issue 更广

把 2026-09-16 至 09-20 之间所有 sandbox:docker 分支拉出来、逐个核对失败日志:2026-09-16 当天有四条分支死于完全相同的 EACCES 签名,跨两台主机——hk1-24 00:27Z、hk2-29 04:26Z、hk2-21 04:39Z、hk2-2 07:45Z(即 #12006 跟踪的那次)。该窗口内其余失败都是别的原因。

从 09-17 04:11Z 起,同样的主机上该分支恢复绿色,也就是说中毒状态在约 21 小时内结束了。是人工清理还是别的原因,日志看不出来——值得一提,因为解析器文件头断言「该状态不会自行愈合:清除它需要有人登录主机」,而生产数据既不能证实也不能证伪这一点。

3. 健康主机上并非完全逐字节一致

没有兄弟分支在跑时,两臂记录到的所有流只差一行新增的 stdout——docker argv、锁位置、退出码全部相同。但当有兄弟分支正在运行时,两臂分叉:labelled prune 都跳过;dangling prune 在 main 上跳过,在 PR 上执行。

PR 把 docker image prune --force --filter until=24h 移出了 daemon 锁的保护,于是它现在会在兄弟分支测试进行中执行。这个选择站得住——dangling 镜像无标签且无引用,有 24h 时限,主机自己的 root 定时器本来就在不持锁的情况下清理它们(qwen-docker-cleanup.sh:44)——但描述里那句「健康主机与旧行为逐字节一致」不准确(中英文两处都有)。

4. 未决 Critical(R7-1)——成立,而且不只是理论

以 root 身份对着一个活着的 leg 运行真实的 qwen-docker-cleanup.sh:

  • 健康主机:leg 持有共享 daemon 锁 → 定时器打印 skipped: sandbox daemon lock busy,labelled prune 从不运行。
  • 中毒主机、PR 臂:leg 在 $RUNNER_TEMP 里,定时器仍解析共享路径,不会因毒文件存在而跳过(:53 只在锁不存在时跳过),以 root 身份无竞争地拿到锁,并在守护进程仍有活跃 leg 的 docker 工作在飞时执行了 labelled prune。

对真实 Docker:一个刚被 leg 通过 docker image inspect 接受的镜像,在定时器的真实命令 until=24h 下存活(阴性对照,刚创建几秒),而在同一命令把阈值移到镜像之后时被删除(阳性对照——移动阈值而不是等 24 小时),随后 docker run 报 No such image。R7-1 的因果链端到端成立。

R7-1 前提只有一半被支持。 它成立的前提是毒物为 runner 属主目录里的 root 属主锁文件,这种情况 runner 无需 sudo 即可 rm -f。两种形态都构造了:runner 属主目录 + root 属主 0400 锁 → rm rc=0,毒物清除,解析器回到共享目录;root 属主目录 + root 属主 0400 锁 → rm rc=1,Permission denied,仍然回退。

而这两种形态在生产中无法区分:heal 的 warning、leg 的 line 46 报错、prune 步骤的 line 2 报错在两者之间逐字节相同。因此 #12006 的日志无法告诉我们 hk2-2 属于哪一种,R7-1 的修复覆盖其中一种,而本 PR 的回退两种都覆盖。

R7-1 提议的修复已实现并实测(只 unlink 探测判定为不可写的锁名,重测,失败则回退):root 属主锁文件的世界里 leg 绿且 daemon 锁留在共享目录、毒物被清成 runner 属主 0644;root 属主目录的世界里正确回退;健康世界不变。它有效、不削弱不可修复的情形,代价是现有 3 个用例(当前断言「锁文件不可写 → 回退」)需要改指到不可修复的形态。那是一个真实的后续项,而不是拦住本 PR 的理由:在它修不了的形态上,回退仍是主机与「分支直接死掉」之间唯一的屏障。

5. 其余未决发现

  • R5-1(未决)——确认。 中毒主机上 prune 步骤在 stderr 打出 ::warning::… using job-private lock dir …,而它的 stdout 说 Docker cleanup skipped,并且它根本没有打开任何锁。同一个步骤里两个互相矛盾的信号。
  • R5-2(未决)——变异确认。 删掉就地回退 else 里的 ci_lock_dir='' 后 56 个用例全绿,而真实主机上的输出从 cannot create a writable <dir>; … needs a human 翻转为 Docker cleanup skipped because the shared daemon is active——一个没有任何探测验证过的成因。该分支按现状的行为是正确的(三个「解析器缺失」世界都实测过),缺的只是覆盖。
  • R9-1——确认。 连回退目录都创建不了时,解析器的 ::error:: 仍会进入注解列表,而步骤本身恢复为 rc=0。
  • R9-2——精确确认。 删掉无条件 dangling prune 会被 2 个用例杀死,但两个都在中毒路径上;健康路径的用例并不钉它。
  • R9-3——确认。 两条 expect(output).not.toMatch(/prune/) 不可能失败:解析器的六条输出模板里没有这个词。
  • R9-4——确认。 被永久禁用的 labelled prune 没有任何注解、也没有运行页摘要,只存在于一行 stdout 里。

除此之外,对 PR 真正保护的东西,变异覆盖是扎实的——8 个变异杀死 7 个,解析器核心的每一个变异都被杀(各 3–9 个用例)。

6. 本次填补的审查缺口

上一轮评审披露了四处缺口,现已填补三处:

  • macOS 泳道:resolve-ci-lock-dir.test.js + e2e-workflow.test.js 在 Darwin 上 56/56(bash 5.3.15),并在 Apple 自带的 bash 3.2.57 下 18/18(/bin/bash,不是 GNU 3.2 容器)。
  • shellcheck(仓库口径,pin 0.11.0):新脚本只产生 1 条 SC2292 style 提示,与既有 run-e2e-tests.sh 已有的约 30 条同类。没有新问题。
  • yamllint(pin 1.35.1,--format github):两臂皆干净。PR 把这一项列为未验证。
  • actionlint(pin 1.7.12):在仓库自己的参数下两臂皆干净(-shellcheck= 关闭了内嵌 shell 检查)。为准确起见:默认参数下 PR 臂会对那句有意为之的 GITHUB_STEP_SUMMARY= bash … 前缀赋值报 SC1007,而 main 静默——这是仓库门禁永远看不到的误报,但「actionlint 已干净解析」只在仓库参数下为真。

建议

合并。 该改动修复的是一个可复现、近期在两台池主机上复发、且在任何测试运行前就让分支死亡的故障;在健康主机上的代价是一行日志和一次主机定时器本来就不持锁执行的 prune。未决 Critical 对中毒主机上的代价描述准确,但在这样的主机上,另一个选项是分支根本跑不起来。

两个值得建 issue 而非拦住本 PR 的后续项:

  1. R7-1 的窄范围修复(上文已实测):它把永久降级变成对 Main CI failed: E2E Tests on aff26a8f36df #12006 triage 所述形态的自愈,代价是改指 3 个用例。
  2. R5-1 / R9-4 的运维信号:中毒主机上运维现在拿到的是一条点名了「步骤从未打开的锁」的警告,而 labelled prune 被永久关闭这件事没有任何注解。

合并前建议更正描述中的两处:「逐字节一致」那句(§3,中英文各一处),以及 Fixes #12006——该 issue 已于 2026-09-18 作为被 #11991 取代而关闭,这条修复应挂在 #11991 下。

wenshao
wenshao previously approved these changes Sep 20, 2026
@wenshao

wenshao commented Sep 20, 2026

Copy link
Copy Markdown
Collaborator

@qwen-code /triage

…ck (#12006)

The resolver vetoed the whole shared dir when a caller-named lock file was
unwritable, but unlink permission lives on the containing directory: a
root-owned lock file inside a runner-writable dir is exactly the #12006
shape, and the runner can clear it without sudo. Unlink only the files the
probe found unwritable, re-test, and fall back only when the repair fails,
so a repairable host keeps cross-job coordination (sibling legs, the root
qwen-docker-cleanup timer, the release-lane build mutex) instead of losing
it until a human logs in.

Co-authored-by: Qwen-Coder <[email protected]>
@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🤖 Addressed the latest review feedback (round 4/10). What changed, and what I pushed back on: · 已处理最新评审反馈(第 4/10 轮)。改动内容与我反驳保留之处如下:

Autofix round — PR #12016 (issue #12006)

Growth audit (required this round)

Verdict: sound — recorded in growth-audit.json before any edit.

  • KISS: pass. Four structurally simpler shapes were considered and each fails the PR's goal: no resolver (that is the Main CI failed: E2E Tests on aff26a8f36df #12006 bug itself), unconditional job-private locks (forfeits coordination on every run, healthy hosts included), tolerating lock-open failure (loses the prune exclusion — the exact timer-vs-live-daemon chain the maintainer reproduced), and the current probe+fallback (minimal coordination-preserving shape). Each accumulated piece is load-bearing for a test-pinned behavior; the test-side overage (410/400) is the bash-driven witness matrix whose gaps were themselves rounds 5–9 findings (R5-2, R9-2, R9-3).
  • Minimal change: pass. All six files trace to Main CI failed: E2E Tests on aff26a8f36df #12006 or an accepted round finding; no untraceable hunks.
  • This round's work is the standing Critical itself; solving it is primary. The repair re-points three existing cases and adds two, keeping growth minimal.

Feedback dispositions

R7-1 (Critical) — rc:4046529395, rc:4049746075, rc:4054490664; standing in reviews rv:5247538768, rv:5251440468, rv:5257402698 — implemented

Claim: the unconditional veto abandons the shared lock dir when a named lock file is unwritable, but in the #12006 shape (root-owned lock file inside a runner-writable dir) the runner can rm -f the poison without sudo, so the job-private fallback trades away cross-job coordination permanently when a repair was possible.

Reproduced on the pre-round code before implementing (probe, this runner, non-root): shared dir writable + docker-sandbox-daemon.lock at 0400 → resolver printed the job-private fallback with ::warning::; a sudo-less rm -f of the same file succeeded; re-running the resolver then printed the shared dir. The claim held exactly as stated.

Fix (.github/scripts/resolve-ci-lock-dir.sh): the veto loop now unlinks only the caller-named files the probe found unwritable (rm -f --, never a glob, never a writable file), re-tests, and falls back only when the lock name still exists and is still unwritable. The header comment's "does not heal itself" premise was corrected to distinguish the repairable shape (root-owned file in a runner-writable dir) from the residual one (unwritable dir, un-unlinkable name), which still needs a human. The resolver's CLI shape, stdout contract, and silent healthy path are unchanged, so the byte-pinned call sites in e2e-workflow.test.js and the daemon-lock path hardcoded by qwen-docker-cleanup.sh:19 and run-release-docker-integration.sh:32 are untouched — the repair is in place, never a relocation.

Tests (scripts/tests/resolve-ci-lock-dir.test.js):

  • New resolver case: writable shared dir + daemon lock at 0400 → resolves to the shared dir, empty stderr, poison unlinked, a fresh lock opens in its place. Red on pre-round code (pre-round behavior = unconditional fallback).
  • New leg case: same fixture driven through the real run-e2e-tests.sh → leg green, daemon lock recreated on the shared path, no job-private fallback. Red on pre-round code.
  • Three existing cases re-pointed from the repairable shape (0400 file in a writable dir) to an unrepairable one: a directory squatting on the lock name at 0500. This is the only single-uid fixture where rm -f fails inside a writable directory (EISDIR), standing in for any unlink failure (EROFS, foreign-owned name in a sticky dir). I deliberately did not use the suggested chmodSync(shared, 0o555) shape for these: an unwritable directory vetoes at the dir check before the lock loop runs, so it never reaches the repair branch and would leave the failed-repair re-test unwitnessed. All original assertions in the three cases were kept (fallback, warning naming the offender, leg green, daemon/build separation); each gained a "poison left untouched" assertion.
  • Narrowness pins added to the two foreign-lock cases: a lock the caller never named is neither repaired nor removed.

Mutation probes (both run, both killed):

  • Drop the repair (restore the unconditional return 1 — i.e. pre-round behavior): the two new repair cases fail, 18 others pass.
  • Drop the post-unlink re-test (trust rm unconditionally): all three re-pointed squat cases fail (the leg cases die at exec 9> on the surviving directory), 17 others pass.
  • (An earlier malformed first mutation that broke the script's syntax failed 15 cases including healthy-path ones; it was discarded and re-applied correctly — only the two valid mutations above are reported as evidence.)

Maintainer verification comment (ic:5747435862, @wenshao)

  • Follow-up pre-release: fix ci #1 — R7-1's narrow repair: implemented in this round rather than deferred, since the loop's three CHANGES_REQUESTED reviews all stand on it and the maintainer's own measurement confirmed the repair works and does not weaken the unrepairable case.
  • Follow-up Where is the config saved? #2 — R5-1/R9-4 operator-signal improvements: remains deferred; both are Suggestion-level and this round is Critical-only under the growth brake.
  • Two PR-description corrections (the "byte-for-byte" sentence in both languages, and retargeting Fixes #12006 to the open Deferred review findings from PR #11974: fix(ci): gate the Linux E2E legs on a disk floor (#11973) #11991): these need a PR-body edit, which this bot cannot perform (no GitHub writes in this mode). Flagging here for the maintainer to apply.

Deferred non-Critical section

Not work — audit record under critical-only mode; no code, thread, or reply actions taken on those items.

Verification

  • npm run build — passed (run as the resumable per-workspace equivalent of scripts/build.js, same dependency order: npm run generate + all 27 workspace build steps + generate-settings-schema; this harness caps foreground commands at 120s and reaps detached processes, so the monorepo build was chunked with per-step markers; every chunk exited 0; the sandbox-image tail is a no-op without BUILD_SANDBOX).
  • npm run typecheck — passed (80s, exit 0).
  • npm run lint — passed (chunked by top-level directory: packages/core, packages/cli, packages/channels, remaining packages, scripts, integrations, integration-tests; every chunk exited 0; neither changed file is in eslint's --ext .ts,.tsx scope).
  • npx prettier --experimental-cli --check on both changed files — passed.
  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/resolve-ci-lock-dir.test.js scripts/tests/e2e-workflow.test.js — 58 passed (20 + 38; resolver suite grew 18 → 20).
  • Full scripts/tests suite — passed via --shard=1/8..8/8 plus the three slow files run individually (install-script 113 passed/16 skipped, acp-serve-boundary-guard 8 passed, qwen-autofix-workflow 338 passed). One failure in scripts/tests/verify-capture.test.js ("preserves colour and bold independently") was reproduced on the unmodified tree (git stash → same failure → git stash pop): pre-existing, environment-related (ANSI rendering), unrelated to this change.
  • Mutation probes — see the R7-1 section above; both mutations killed, suite green after restore.
  • shellcheck — not installed in this environment (the maintainer's pinned 0.11.0 run found only the pre-existing SC2292 style class; this change adds no new shell constructs beyond rm -f -- … || true and an [[ -e && ! -w ]] re-test).

Notes

  • No test weakening: no file deleted, no test disabled, no assertion removed; enabled resolver-suite tests went 18 → 20 and the three re-pointed cases kept every assertion they had (each gained one). No test-weakening.json required.
  • No comment-replies.json: every inline finding this round (the three R7-1 threads) is resolved in code.
  • No base-conflict resolution requested (--conflict false); no merge performed.
中文说明

Autofix 本轮结果 — PR #12016(issue #12006)

增长审计(本轮必需)

结论:sound——已在任何编辑之前写入 growth-audit.json。

  • KISS:通过。 评估过四种结构上更简单的形态,均无法达成 PR 目标:不加 resolver(那就是 Main CI failed: E2E Tests on aff26a8f36df #12006 故障本身)、无条件使用作业私有锁(健康主机也永久失去协调)、容忍锁打开失败(失去 prune 互斥——正是 maintainer 实测过的「定时器 vs 存活 daemon」链),以及现有的探测+回退(保持协调的最小形态)。现有每一部分都有测试钉住的负载理由;测试侧超支(410/400)来自 bash 驱动的见证矩阵,而第 5–9 轮的发现(R5-2、R9-2、R9-3)本身就是覆盖缺口。
  • 最小改动:通过。 六个文件全部可追溯到 Main CI failed: E2E Tests on aff26a8f36df #12006 或已接受的评审发现;无来源不明的 hunk。
  • 本轮工作就是未决 Critical 本身;解决它是首要的。修复改指了三个既有用例、新增两个,增长控制在最小。

反馈处理

R7-1(Critical)—— rc:4046529395、rc:4049746075、rc:4054490664;并立于评审 rv:5247538768、rv:5251440468、rv:5257402698 —— 已实现

主张:当一个被点名的锁文件不可写时,无条件否决放弃了整个共享锁目录;但在 #12006 的形态下(runner 可写目录里的 root 属主锁文件),runner 无需 sudo 即可 rm -f 清除毒物,因此作业私有回退在本可修复的情况下永久牺牲了跨作业协调。

实现前已在改动前代码上复现(本机非 root 探测):共享目录可写 + docker-sandbox-daemon.lock 为 0400 → resolver 输出作业私有回退并打 ::warning::;同一文件无需 sudo 的 rm -f 成功;再次运行 resolver 即输出共享目录。主张完全成立。

修复(.github/scripts/resolve-ci-lock-dir.sh):否决循环现在只 unlink 探测判定为不可写的、由调用方点名的文件(rm -f --,绝不通配,绝不动可写文件),随后重新检测;仅当该锁名仍存在且仍不可写时才回退。文件头注释中「不会自愈」的前提已更正,以区分可修复形态(runner 可写目录里的 root 属主文件)与仍需人工的残余形态(目录不可写、名字无法 unlink)。resolver 的 CLI 形态、stdout 契约与健康路径的静默均未改变,因此 e2e-workflow.test.js 逐字节锚定的调用点、以及 qwen-docker-cleanup.sh:19 与 run-release-docker-integration.sh:32 硬编码的 daemon 锁路径都不受影响——修复是就地进行的,绝不挪移。

测试(scripts/tests/resolve-ci-lock-dir.test.js):

  • 新增 resolver 用例:可写共享目录 + 0400 daemon 锁 → 解析到共享目录、stderr 为空、毒物被 unlink、原位置可新建锁。在改动前代码上为红(改动前行为 = 无条件回退)。
  • 新增 leg 用例:同一 fixture 驱动真实的 run-e2e-tests.sh → leg 绿、daemon 锁在共享路径上重建、无作业私有回退。在改动前代码上为红。
  • 三个既存用例从可修复形态(可写目录中的 0400 文件)改指到不可修复形态:一个 0500 目录占据锁名。这是单 uid 测试中唯一能让 rm -f 在可写目录内失败的 fixture(EISDIR),代表一切 unlink 失败(EROFS、sticky 目录中的他人文件)。我刻意没有采用建议的 chmodSync(shared, 0o555) 形态:目录不可写会在目录检查处先否决,根本走不到锁循环,修复失败后的重测分支将无人见证。三个用例的全部原有断言都保留(回退、警告点名肇事者、leg 绿、daemon/build 分离),并各新增一条「毒物保持原样」断言。
  • 两个外部锁用例新增窄化断言:调用方未点名的锁既不被修复也不被删除。

变异探测(两个均已执行、均被杀死):

  • 去掉修复(恢复无条件 return 1,即改动前行为):两个新增修复用例失败,其余 18 个通过。
  • 去掉 unlink 后的重测(无条件相信 rm):三个改指后的 squat 用例全部失败(leg 用例在残留的目录上死于 exec 9>),其余 17 个通过。
  • (早先一次写坏的变异破坏了脚本语法,导致含健康路径在内的 15 个用例失败;该次已废弃并正确重做——只有上述两个有效变异作为证据报告。)

Maintainer 验证评论(ic:5747435862,@wenshao)

  • 后续项 1——R7-1 的窄修复:本轮已实现而非延后,因为循环里三条 CHANGES_REQUESTED 评审都立在它上面,且 maintainer 自己的实测确认该修复有效且不削弱不可修复情形。
  • 后续项 2——R5-1/R9-4 运维信号改进:保持延后;两者均为 Suggestion 级,而本轮在增长制动下仅处理 Critical。
  • 两处 PR 描述更正(两种语言的「逐字节一致」句,以及把 Fixes #12006 改挂到未关闭的 Deferred review findings from PR #11974: fix(ci): gate the Linux E2E legs on a disk floor (#11973) #11991):需要编辑 PR 正文,本 bot 在此模式下无 GitHub 写权限,无法执行。在此标记给 maintainer 处理。

延后的非 Critical 区段

非工作项——仅处理 Critical 模式下的审计记录;未对这些条目做代码、线程或回复动作。

验证

  • npm run build——通过(以 scripts/build.js 的可续跑分块等价形式执行,依赖顺序相同:npm run generate + 全部 27 个 workspace 构建步骤 + generate-settings-schema;本环境前台命令上限 120s 且会回收脱离进程,因此 monorepo 构建按分块标记逐段完成;每段退出码均为 0;未设置 BUILD_SANDBOX 时沙箱镜像尾部为空操作)。
  • npm run typecheck——通过(80s,退出码 0)。
  • npm run lint——通过(按顶层目录分块:packages/core、packages/cli、packages/channels、其余 packages、scripts、integrations、integration-tests;每块退出码均为 0;两个改动文件都不在 eslint 的 --ext .ts,.tsx 范围内)。
  • npx prettier --experimental-cli --check 检查两个改动文件——通过。
  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/resolve-ci-lock-dir.test.js scripts/tests/e2e-workflow.test.js——58 通过(20 + 38;resolver 套件从 18 增至 20)。
  • 完整 scripts/tests 套件——通过,经 --shard=1/8..8/8 加三个慢文件单独运行(install-script 113 通过/16 跳过,acp-serve-boundary-guard 8 通过,qwen-autofix-workflow 338 通过)。scripts/tests/verify-capture.test.js 中一处失败(「preserves colour and bold independently」)已在未改动的树上复现(git stash → 同样失败 → git stash pop):既存问题,与环境相关(ANSI 渲染),与本改动无关。
  • 变异探测——见上文 R7-1 一节;两个变异均被杀死,恢复后套件全绿。
  • shellcheck——本环境未安装(maintainer 以 pin 0.11.0 的运行只发现既有的 SC2292 style 类;本改动除 rm -f -- … || true 与 [[ -e && ! -w ]] 重测外未引入新的 shell 结构)。

说明

  • 无测试削弱:未删文件、未禁用用例、未删断言;resolver 套件的启用用例从 18 增至 20,三个改指用例保留了原有全部断言(并各增一条)。无需 test-weakening.json。
  • 无 comment-replies.json:本轮全部行内发现(三条 R7-1 线程)已在代码中解决。
  • 未请求处理基线冲突(--conflict false);未做合并。

Deferred non-Critical feedback

Critical-only mode is active: the PR's diff grew src 33 / test 410 net lines beyond this counting window's baseline (budgets: 400/400). The workflow excluded the non-Critical feedback below from this round's actionable sections; the items remain open for human follow-up. Maintainer feedback continues to flow unaffected during a growth-only engagement (the per-author batch budget applies only after 5 change-producing rounds). (@qwen-code /retry starts a fresh counting window.)

中文说明

已进入仅处理 Critical 的模式:本计数窗口内 diff 净增长已达 源码 33 / 测试 410 行(预算 400/400)。本轮可执行区域已排除下方非 Critical 反馈;这些条目保持开放,留待人工跟进。纯增长触发期间维护者反馈照常流动(按作者的批次预算仅在完成 5 个产生改动的轮次后生效)。(评论 @qwen-code /retry 可开启新的计数窗口。)

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。

🧵 Resolved all 1 selected review thread(s). · 已关闭全部选中的 1 条评审线程。

Re-review when you have a moment. After round 10 this bot stops and leaves the PR for a human. · 有空请复审;第 10 轮后本 bot 停止并将 PR 交给人工。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.24.1

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🔀 Base updated: red check(s) [Test (ubuntu-latest, Node 22.x), Lint & Static (ubuntu-latest, Node 22.x), Integration Tests (no-AK, No Sandbox), web-shell E2E Smoke (ubuntu-latest, Node 22.x)] pass on current main — merged current main via update-branch; CI will re-run.

中文说明

🔀 已更新 base:红色检查 [Test (ubuntu-latest, Node 22.x), Lint & Static (ubuntu-latest, Node 22.x), Integration Tests (no-AK, No Sandbox), web-shell E2E Smoke (ubuntu-latest, Node 22.x)] 在当前 main 上通过 —— 已通过 update-branch 合入当前 main,CI 将重新运行。

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🤖 AutoFix updated a stale base — the fix did not pass verification, but this PR was behind main, so it merged current main in via update-branch and will retry on the next scan. A stale base (a dependency or symbol main already changed) can fail the build without being the fix's fault; if it still fails once current, it hands off to a human.

What I found before stopping:
Qwen hit the tool-call loop guard during address-review. A human should take over this feedback batch.

See the Qwen Autofix agent step logs for model/tool output.

中文说明

🤖 AutoFix 更新了一个过期的 base —— 修复未通过验证,但本 PR 落后于 main,因此已通过 update-branch 合入当前 main,并将在下次扫描时重试。过期的 base(main 已改动的依赖或符号)可能让构建失败而并非修复本身的错;若 base 更新后仍然失败,将移交人工处理。

Run log: https://github.com/QwenLM/qwen-code/actions/runs/35497432176


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.24.1

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🤖 Reviewed the latest feedback — no changes needed. Why, point by point: · 已审阅最新反馈——无需改动。逐点说明原因如下:

Autofix round — PR #12016 (issue #12006): no code change this round

Why no change

Every actionable finding is already resolved in code at the current head (1effcfbedc), and the maintainer has explicitly steered this PR to merge with the remaining items filed as follow-ups.

  • The one standing Critical (R7-1) was fixed by e2993f3e76 ("repair an unwritable CI lock file in place before falling back"), which landed after the latest review (round 9, reviewed 1cf3bc95). The resolver now unlinks only the caller-named lock files its probe found unwritable, re-tests, and falls back to the job-private dir only when the repair fails — exactly the narrow repair the maintainer independently implemented and measured. I re-verified it at the current head; see Verification below.
  • The maintainer's verification (@wenshao, issue comment 5747435862) recommends merge. Follow-up pre-release: fix ci #1 (R7-1's narrow repair) is done in e2993f3e76; follow-up Where is the config saved? #2 (R5-1/R9-4 diagnostics) and the remaining Suggestions are explicitly "worth filing rather than blocking on". This PR has been through ~9 review rounds, so per the review-rounds policy it now lands Criticals only and defers Suggestions.
  • CI is green at this exact head — all required checks SUCCESS; macOS/Windows lanes are trigger-gated off pull requests by design.

Resolved this round (re-verified, fixed by an earlier commit)

  • R7-1 (rc:4046529395, rc:4049746075, rc:4054490664): repair-in-place confirmed working at 1effcfbedc — see Verification. Threads listed in resolved-comments.txt.

Deferred to the follow-up queue (recorded, threads left open with a reply)

Per the maintainer's recommendation and the convergence posture, no Suggestion-level item is implemented in this PR:

  • rc:4036007801 (R5-1) and R9-4 — the maintainer's follow-up Where is the config saved? #2 (operator diagnostics on a poisoned host).
  • rc:4036007811 (R5-2) — coverage of the inline fallback's else arm (maintainer confirmed the gap by mutation).
  • rc:4025497792 (R1-2 remainder) — resolver adoption at the release lane / sdk-java.yml and a host-shared fallback tier: outside this PR's footprint.
  • rc:4031664803 (R4-1/R7-2), rc:4031664817 (R4-2), rc:4031664822 (R4-3) — diagnostics accuracy / coverage, same follow-up cluster.
  • R9-1, R9-2, R9-3 (round-9 deferrals, review 5257402698) — maintainer-confirmed diagnostics/coverage gaps.

All recorded in deferred-findings.json; the inline threads have bilingual replies in comment-replies.json.

Needs the maintainer (this bot has no GitHub writes)

Two PR-description corrections from the verification comment remain open — they require editing the PR body, which only a human or the PR workflow can do:

  1. The "byte-for-byte" sentence (§3 of the verification, both languages): with a sibling leg live, the PR runs the dangling prune that main skips.
  2. Fixes #12006 → retarget to Deferred review findings from PR #11974: fix(ci): gate the Linux E2E legs on a disk floor (#11973) #11991 (issue Main CI failed: E2E Tests on aff26a8f36df #12006 was closed on 2026-09-18 as superseded).

Note on the two triage-verify failures

Issue comments 5747219213 and 5747495869 report the sandboxed-verification harness failing at pnpm install before any verification started. This is not attributable to the PR's content: the diff touches no dependency manifest or lockfile (only .github/scripts/*.sh, .github/workflows/e2e.yml, and scripts/tests/*), a clean install was already performed on this runner before this round, the focused suites execute green locally, and all PR CI checks pass at this head.

Verification

  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/resolve-ci-lock-dir.test.js scripts/tests/e2e-workflow.test.js — 58 passed (20 + 38), 0 failed.
  • Behavioral probe of R7-1's fix at this head (real bash, non-root, temp HOME): (a) writable shared dir + docker-sandbox-daemon.lock at 0400 → resolver prints the shared dir, rc=0, poison unlinked; (b) a directory squatting on the lock name at 0500 → resolver falls back to the job-private dir with ::warning:: naming the offender, rc=0, poison untouched.
  • git status — clean; HEAD 1effcfbedc (matches origin). All 14 checks at this SHA: SUCCESS/SKIPPED, none failing.

No commit was created; the working tree is unchanged.

中文说明

Autofix 本轮结论 —— PR #12016(issue #12006):本轮无代码改动

为什么不做改动

在当前 head(1effcfbedc)上,所有需要处理的发现都已在代码中解决,且维护者已明确指示本 PR 走向合并、其余事项以跟进项建档。

  • 唯一未决的 Critical(R7-1)已由 e2993f3e76 修复("repair an unwritable CI lock file in place before falling back"),该提交晚于最近一次评审(第 9 轮,评审对象为 1cf3bc95)。resolver 现在只 unlink 探测判定为不可写的、由调用方点名的锁文件,随后重新探测,仅在修复失败时才回退到作业私有目录——这正是维护者独立实现并实测过的窄范围修复。我已在当前 head 上重新验证,见下方「验证」。
  • 维护者的验证结论(@wenshao,issue 评论 5747435862)建议合并。 跟进项 pre-release: fix ci #1(R7-1 的窄范围修复)已在 e2993f3e76 中完成;跟进项 Where is the config saved? #2(R5-1/R9-4 诊断信息)及其余建议级条目被明确列为「值得建档跟进,而非阻塞合并」。本 PR 已历经约 9 轮评审,按评审轮次策略,此后只落地 Critical 级修复,建议级条目一律转入跟进。
  • 当前 head 的 CI 全绿——所有必需检查均为 SUCCESS;macOS/Windows 泳道按设计不在 pull request 上触发。

本轮解决(重新验证,由先前提交修复)

  • R7-1(rc:4046529395、rc:4049746075、rc:4054490664):在 1effcfbedc 上确认原位修复生效——见「验证」。相关线程已列入 resolved-comments.txt。

转入后续跟进队列(已记录,线程回复后保持未解决)

按维护者建议与收敛姿态,本 PR 不再实现任何建议级条目:

  • rc:4036007801(R5-1) 与 R9-4——维护者指定的跟进项 Where is the config saved? #2(中毒主机上的运维诊断信息)。
  • rc:4036007811(R5-2)——就地回退 else 分支的测试覆盖(维护者已用变异测试确认该缺口)。
  • rc:4025497792(R1-2 剩余部分)——在 release 泳道 / sdk-java.yml 采用 resolver,以及主机共享回退层级:超出本 PR 范围。
  • rc:4031664803(R4-1/R7-2)、rc:4031664817(R4-2)、rc:4031664822(R4-3)——诊断准确性 / 覆盖,属同一跟进簇。
  • R9-1、R9-2、R9-3(评审 5257402698 的第 9 轮延后项)——维护者确认过的诊断/覆盖缺口。

以上全部记录于 deferred-findings.json;相关行内线程的中英文回复见 comment-replies.json。

需要维护者处理(本 bot 无 GitHub 写权限)

验证评论中要求的两处 PR 描述更正仍未完成——它们需要编辑 PR 正文,只能由人工或 PR 工作流执行:

  1. 「逐字节一致」那句(验证报告 §3,中英文各一处):当有兄弟分支在跑时,本 PR 会执行 main 所跳过的 dangling prune。
  2. Fixes #12006 应改挂 Deferred review findings from PR #11974: fix(ci): gate the Linux E2E legs on a disk floor (#11973) #11991(issue Main CI failed: E2E Tests on aff26a8f36df #12006 已于 2026-09-18 作为被取代而关闭)。

关于两次 triage-verify 失败的说明

issue 评论 5747219213 与 5747495869 报告沙箱验证装置在任何验证开始前的 pnpm install 阶段失败。这不能归因于本 PR 的内容:diff 未触碰任何依赖清单或 lockfile(只改了 .github/scripts/*.sh、.github/workflows/e2e.yml 和 scripts/tests/*);本轮开始前本 runner 已完成过一次干净安装;聚焦测试套件本地执行全绿;且当前 head 上 PR 的全部 CI 检查均通过。

验证

  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/resolve-ci-lock-dir.test.js scripts/tests/e2e-workflow.test.js——58 通过(20 + 38),0 失败。
  • 在当前 head 上对 R7-1 修复的行为级探针(真实 bash、非 root、临时 HOME):(a) 共享目录可写 + docker-sandbox-daemon.lock 为 0400 → resolver 输出共享目录,rc=0,毒文件被 unlink;(b) 目录占用锁文件名(0500)→ resolver 回退到作业私有目录,::warning:: 点名出错文件,rc=0,毒物保持不动。
  • git status——干净;HEAD 为 1effcfbedc(与 origin 一致)。该 SHA 上 14 项检查:SUCCESS/SKIPPED,无失败。

本轮未创建提交;工作树未改动。

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。

🧵 Resolved all 1 selected review thread(s). · 已关闭全部选中的 1 条评审线程。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.24.1

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🕐 Review received — an automatic review of the current head is still running, so this round is held until it lands (a push now would cancel it and discard its work, #8888). Your feedback stays queued for the next eligible round.

中文说明

🕐 已收到评审 —— 当前 head 上仍有一轮自动 review 在运行,本轮暂缓(现在推送会取消该 review 并丢弃其工作,#8888)。反馈保持排队,等待下一次可运行的轮次处理。

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🤖 Addressed the latest review feedback (round 2/10). What changed, and what I pushed back on: · 已处理最新评审反馈(第 2/10 轮)。改动内容与我反驳保留之处如下:

Autofix round — PR #12016 (issue #12006)

Commit: 301c1bb7d1 fix(ci): never unlink a live-held CI lock file in the resolver (#12006)

Feedback addressed

[rc:4057016408] — Critical R7-1: the in-place heal unlinks a lock file on the sole evidence of a permission bit — implemented

Reproduced first (source-blind). With a live holder mirroring the root qwen-docker-cleanup timer (exec 8>> + flock --exclusive + chmod 0400) on docker-sandbox-daemon.lock, the pre-fix resolver: exited 0 certifying the shared dir, printed no ::warning::, and unlinked the held inode (ino_after=GONE). Probe script and output are reproducible from this round's log.

Fix (.github/scripts/resolve-ci-lock-dir.sh): a new lock_held probe runs before any rm -f; a lock that cannot be proved unheld is refused, and the resolver falls back to the job-private dir with the cause named (... is not writable and is locked by a live process). Two probes, because each sees a shape the other cannot:

  1. a read-only open plus a non-blocking exclusive flock — refused exactly when a holder sits on the inode, and < can never truncate a held one (the >>-never-> constraint is honoured by never opening for write at all);
  2. /proc/locks (read as the job uid, no access to the file needed) for the finding's exact production shape — the root-owned 0400 daemon lock that cannot even be opened. Where neither probe exists (macOS has neither flock(1) nor /proc/locks) the probe reports not-held, leaving those lanes on the previous behaviour.

Also per the finding: the successful heal now emits a ::warning:: on stderr naming the file it unlinked (stdout stays the single-path payload both callers capture with $(...)), and the stale comment that claimed the re-test covers a root process holding the old inode was corrected (the re-test never saw that case — after a successful rm -f, -e is false). The daemon lock is repaired in place or refused — never relocated.

Tests (scripts/tests/resolve-ci-lock-dir.test.js):

  • New case refuses to unlink a lock a live process holds: a holder takes flock -x on the shared daemon lock and sleeps; asserts the resolver prints the job-private fallback with a ::warning::, statSync(daemonLock).ino is unchanged, and a contender flock --nonblock -x fails while the holder sleeps. Skips only where the fixture cannot exist (root, or no flock(1) — macOS).
  • Updated repairs an unwritable lock file in place...: the heal now pins the ::warning:: naming the file instead of pinning an empty stderr. An unheld foreign-owned 0400 file is still unlinked, as the finding requires.

Mutation probes (guard deleted → test red → restored → green):

  • lock_held guard removed → refuses to unlink a lock a live process holds FAILED (×), 20 others green; restored → 21/21 green.
  • Heal ::warning:: removed → repairs an unwritable lock file in place... FAILED (×); restored → 21/21 green.

Both changed tests therefore fail on the pre-round code and pass on the round's code.

Not addressed (out of scope this round)

  • The four R10-* items and the eight re-deferred Suggestions were explicitly listed as "recorded, not requested in this round" / already deferred in rounds 5–9. Left untouched.
  • verify-capture.test.js > preserves colour and bold independently fails in the full scripts run. It fails identically with this round's two files restored to HEAD (verified: snapshot diff → git checkout -- → rerun → same failure → re-applied), so it is a pre-existing, environment-caused failure (bold font rendering in this container), not a regression from this round. It lives far outside this PR's footprint.

Verification

  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/resolve-ci-lock-dir.test.js — 21 passed (also green in the full-suite run: ✓ ... (21 tests))
  • npm run test:scripts — 2566 passed, 19 skipped, 1 failed; the single failure is the pre-existing verify-capture bold-font case proven above to reproduce at base content
  • npm run build — passed (exit 0)
  • npm run typecheck — passed (exit 0)
  • npm run lint — passed (exit 0)
  • Live reproduction probe (holder + real resolver, before/after) — before: held lock unlinked, shared dir certified, stderr silent; after: job-private fallback with ::warning::, inode unchanged, contender flock refused
  • Mutation probes — both new witnesses go red when their guard is removed (see above)
  • Integration tests via npm run bundle — not run: the touched behavior (CI workflow shell scripts) is never exercised through the bundled CLI; the bash-driven scripts suite above is its harness
  • Note: /proc/locks is blind in this sandbox (0 entries while a lock is verifiably held), so the /proc/locks branch was verified by inspection here; the flock-probe branch — the one the new test drives — was executed for real. On real-kernel CI runners /proc/locks covers the root-owned 0400 case from the finding.
中文说明

Autofix 本轮报告 — PR #12016(issue #12006)

提交:301c1bb7d1 fix(ci): never unlink a live-held CI lock file in the resolver (#12006)

已处理的反馈

[rc:4057016408] — Critical R7-1:就地治愈仅凭权限位就 unlink 锁文件 —— 已实现

先复现(source-blind)。 用一个完全复刻 root qwen-docker-cleanup 定时器的活持有者(exec 8>> + flock --exclusive + chmod 0400)持有 docker-sandbox-daemon.lock 时,修复前的解析器:退出码 0 并判定共享目录健康、没有 ::warning::、且 unlink 了被持有的 inode(ino_after=GONE)。

修复(.github/scripts/resolve-ci-lock-dir.sh):新增 lock_held 探测,在任何 rm -f 之前运行;无法证明无人持有的锁一律拒绝,解析器回退到作业私有目录并点名原因(... is not writable and is locked by a live process)。两个探测互补,各自覆盖对方看不到的形态:

  1. 只读打开加非阻塞独占 flock——有活持有者时必定被拒绝,且 < 永远不可能截断被持有的 inode(通过完全不以写方式打开来满足「用 >> 绝不用 >」的约束);
  2. /proc/locks(以作业 uid 读取,不需要访问文件本身)——覆盖本条发现的确切生产形态:连打开都打不开的 root 属主 0400 daemon 锁。两个探测都不可用时(macOS 既没有 flock(1) 也没有 /proc/locks)按未持有处理,这些泳道保持原有行为。

同时按发现要求:治愈成功后在 stderr 发出点名被 unlink 文件的 ::warning::(stdout 保持为两个调用方用 $(...) 捕获的单一路径负载);并修正了那条声称「重测覆盖了 root 进程持有旧 inode」的陈旧注释(重测根本看不到这种情形——rm -f 成功后 -e 为假)。daemon 锁只在原地修复或拒绝——绝不挪走。

测试(scripts/tests/resolve-ci-lock-dir.test.js):

  • 新用例 refuses to unlink a lock a live process holds:持有者对共享 daemon 锁取 flock -x 后 sleep;断言解析器输出带 ::warning:: 的作业私有回退、statSync(daemonLock).ino 不变、持有期间竞争者 flock --nonblock -x 失败。仅在夹具无法存在的环境跳过(root,或没有 flock(1)——macOS)。
  • 更新 repairs an unwritable lock file in place...:治愈路径从断言 stderr 为空改为断言点名该文件的 ::warning::。按发现的要求,无人持有的外来 0400 文件仍会被 unlink。

变异探针(删除守卫 → 用例转红 → 恢复 → 转绿):

  • 删除 lock_held 守卫 → refuses to unlink a lock a live process holds 失败(×),其余 20 条绿;恢复后 21/21 绿。
  • 删除治愈 ::warning:: → repairs an unwritable lock file in place... 失败(×);恢复后 21/21 绿。

因此两个改动过的用例在本轮之前的代码上都会失败、在本轮代码上通过。

本轮未处理(超出范围)

  • 四条 R10-* 与八条此前已延后的 Suggestion 在反馈中被明确标注为「已记录,本轮不要求修改」/ 第 5–9 轮已延后。未触碰。
  • verify-capture.test.js > preserves colour and bold independently 在完整脚本套件中失败。已验证该失败与本轮无关:把本轮的两个文件恢复到 HEAD 后失败完全相同(快照 diff → git checkout -- → 重跑 → 同一失败 → 重新应用),属于本容器环境(粗体字体渲染)导致的既有失败,且远在本 PR 的改动范围之外。

验证

  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/resolve-ci-lock-dir.test.js — 21 通过(完整套件中同样通过:✓ ... (21 tests))
  • npm run test:scripts — 2566 通过、19 跳过、1 失败;唯一失败即上文已证明在基线同样复现的 verify-capture 粗体字体用例
  • npm run build — 通过(exit 0)
  • npm run typecheck — 通过(exit 0)
  • npm run lint — 通过(exit 0)
  • 活持有者对真实解析器的前后复现探针 —— 之前:被持有的锁被 unlink、共享目录被判健康、stderr 静默;之后:带 ::warning:: 的作业私有回退、inode 不变、竞争者 flock 被拒绝
  • 变异探针 —— 两个新见证在删除各自守卫时均转红(见上)
  • npm run bundle 后的集成测试 —— 未运行:本轮改动的行为(CI workflow shell 脚本)不经过打包后的 CLI,上面的 bash 驱动脚本套件即其测试线束
  • 说明:本沙箱的 /proc/locks 是盲的(锁被确实持有时仍 0 条记录),故 /proc/locks 分支在此仅经走查验证;flock 探测分支——即新测试驱动的分支——是真实执行的。在真实内核的 CI runner 上,/proc/locks 覆盖发现中 root 属主 0400 的情形。

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。

🧵 Resolved all 1 selected review thread(s). · 已关闭全部选中的 1 条评审线程。

Re-review when you have a moment. After round 10 this bot stops and leaves the PR for a human. · 有空请复审;第 10 轮后本 bot 停止并将 PR 交给人工。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.24.1

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator Author

🤖 Addressed the latest review feedback (round 3/10). What changed, and what I pushed back on: · 已处理最新评审反馈(第 3/10 轮)。改动内容与我反驳保留之处如下:

Round summary — PR #12016 (issue #12006)

Commit: 10540b602b fix(ci): correct the /proc/locks witness and vet lock openability in the resolver (#12006)

Both Critical findings from round 11 were reproduced and fixed at the root, each with the test coverage the finding asked for. No feedback was declined, deferred, or escalated. No merge conflicts (--conflict false).

[rc:4058048832] Critical R7-1 — /proc/locks witness keyed on the PID field — RESOLVED

Reproduced. A /proc/locks line is <id>: FLOCK ADVISORY WRITE <pid> <major>:<minor>:<inode> <start> <end>: $5 is the holder's PID, $6 the device:inode triple (verified against /proc/self/fdinfo/8 on a live flock). The shipped awk split $5, so a[3] was always empty and the witness could never report "held" — inert exactly on the unreadable root-owned-lock path (flock probe exits 66) that the guard exists for. End-to-end probe of the shipped script against a real live holder on a 0000 lock (flock probe genuinely exiting 66, witness line fed through an awk shim because this sandbox masks /proc/locks): the resolver certified the shared dir, printed ::warning::unlinked stale unwritable lock file …, and unlinked the live-held inode (stat after: inode gone).

Fix. lock_held now builds major:minor:inode from stat -c %D/%i (hex arithmetic in bash, not awk — strtonum is gawk-only) and matches $6 against it, comparing the device with the inode so an inode-only match cannot name a phantom holder on another filesystem. The daemon lock path, the resolver's CLI shape, and the heal's >>-only probing are unchanged, per the finding's constraints. Same probe after the fix: job-private fallback, is not writable and is locked by a live process on stderr, held inode preserved.

Tests. Added beside refuses to unlink a lock a live process holds in scripts/tests/resolve-ci-lock-dir.test.js:

  • refuses to unlink an unreadable lock a live process holds — the requested 0000 case (identical exec 8>> + exclusive flock holder, gated isRoot || !hasFlock || !hasProcLocksWitness), asserting the job-private fallback on stdout, the held-warning on stderr, the inode unchanged, and a contender's flock --nonblock --exclusive still failing. One deliberate deviation from the requested assertions: the contender probe runs after restoring the mode to 0600, because a 0000 file refuses every open with EACCES and the probe would never reach the flock it is meant to measure. The gate also carries one deliberate addition: hasProcLocksWitness (a hold-a-flock-and-look capability probe) replaces a bare existsSync('/proc/locks'), because sandboxes can present /proc/locks as an empty masked file — as this runner does — where the witness environment genuinely cannot exist and the case would assert a veto nothing can produce. On normal Linux CI (and the pool runners the script targets) the probe passes and the case runs.
  • heals an unreadable lock no live process holds — the requested negative sibling (unheld 0000 lock must still be healed), which is what catches an always-held witness mutant.

Mutation probes.

  • Always-held witness mutant (a return 0 planted after the awk): the negative sibling goes RED (Tests 1 failed | 24 skipped), restored byte-identical after.
  • $6→$5 revert: the committed positive case capability-skips on this masked-/proc/locks sandbox, so the revert was instead driven end-to-end through the real script with the awk shim: the $5 mutant unlinks the live-held 0000 lock and certifies the shared dir (inode gone), while the fixed script vetoes and preserves it — the exact properties the committed test asserts wherever /proc/locks works, where it goes red on the revert.

[rc:4058048841] Critical R11-1 — mode-bit gate where callers then open() — RESOLVED

Reproduced. The two requested prune-step cases were written first and run against the pre-round code: both RED. With a runner-owned 0755 directory at the daemon-lock name the step printed Docker cleanup skipped because the shared daemon is active (the EISDIR open failure swallowed by the elif condition under bash -e, misreported as a live daemon); with a dangling symlink whose target sits inside the same writable directory the labelled docker image prune --all --force ran behind a lock that excludes nothing.

Fix (one structural close, not per-entrance patches). In resolve-ci-lock-dir.sh's probe loop, before the mode-bit branch: refuse any name that exists (or is a dangling symlink — -e follows links, so -L sees it) but is not a regular file; after the existing heal: vet the very open the caller makes next with (exec 9>>…) — append mode so the probe can never truncate an inode a live process flocks — gated on the name still existing so a just-healed name is not re-created ahead of the caller's own open (keeps the pinned existsSync(...) === false heal witness green). The regular-file vet is what keeps that open bounded (a FIFO would hang it). In the e2e.yml prune step the exec 9> moved out of the elif condition into its own if ! exec … arm with an explicit cannot open … for writing skip line, so bash -e can no longer fold an open failure into the shared daemon is active message; the resolver-fallback skip reason now names the daemon lock path. The byte-for-byte pin in scripts/tests/e2e-workflow.test.js was updated to the new shape (if ! exec 9>… + elif flock --nonblock 9), assertion count unchanged (one pin became two). e2e.yml stays within its size ratchet (35969 B vs 33227 B baseline + 4096 B allowance).

Tests. Added under describe('prune step daemon lock discipline'), both driving runPruneStep(world, { withResolver: true }) and both red pre-round:

  • skips the labelled prune when a directory squats on the daemon lock name (default 0755, not the 0500 the existing cases use), and
  • skips the labelled prune when a dangling symlink squats on the daemon lock name (target inside the same writable directory),

each asserting no shared daemon is active, a skip line naming docker-sandbox-daemon.lock, no image prune --all in the docker log, and the unconditional dangling prune still running; the symlink case also pins that the squatter is refused, never repaired away.

Mutation probe. Reverting the resolver's vetting to the bare [[ -e && ! -w ]] gate turns the symlink case RED. The directory case stays green under that single revert because the step-side open arm — the finding's other requested change — independently catches EISDIR with a truthful skip line; against the fully pre-round code both cases are red (verified directly).

Notes for reviewers

  • The six [probe] items in the review's deferred section were recorded there by the reviewer as non-blocking for this round and were not worked. One of them (the inode-only phantom-holder gap) is closed as a load-bearing part of the R7-1 fix; the rest (the lock_held check-then-act race, the /proc/locks-branch coverage, the banner's permanent-human framing, the leg-side labelled prune, the strict-prefix pin) remain deferred as recorded.
  • No CI machinery outside this PR's own footprint was touched; all four changed files were already in the PR diff. No test was deleted, disabled, or weakened — one pin string was updated to the new prune-step shape and its assertion count grew.

Verification

  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/resolve-ci-lock-dir.test.js — 24 passed, 1 skipped (the 0000 live-holder case capability-skips on this sandbox's masked /proc/locks; runs on normal Linux CI)
  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/e2e-workflow.test.js — 38 passed
  • npx vitest run … scripts/tests/workflow-size.test.js scripts/tests/no-ak-integration-ci.test.js (same config, with the two files above) — 4 files, 290 passed, 1 skipped
  • New prune-step cases against the pre-round code — both failed as reproduced defects (red-first)
  • Mutation probes — always-held witness: negative sibling red; resolver-vetting revert: symlink case red; $5 revert via awk shim against a real live holder: live-held inode unlinked by the mutant, preserved by the fix (details above)
  • bash -n .github/scripts/resolve-ci-lock-dir.sh — syntax ok
  • npx prettier --check on all four changed files — clean (after --write on the test file; suites re-run green after formatting)
  • npm run typecheck — passed
  • npm run build — passed; run as the exact per-package sequence of scripts/build.js (generate → browser-use → core → channels → audio-capture → node-repl → acp-bridge → sdk-typescript → web-shell → web-templates → cli → settings-schema → qwen-live → vscode-ide-companion → chrome-extension → external-context ×2), every step rc=0, because this runner kills any single command at 120s and the monolithic invocation exceeds it; the opt-in sandbox image build (BUILD_SANDBOX=1) was not run, and generate-settings-schema produced no diff
  • npm run lint (eslint . --ext .ts,.tsx + eslint integration-tests) — passed; run in chunks under the same 120s limit (scripts + integration-tests, packages/cli, packages/core, all remaining packages/integrations, root-level files), every chunk rc=0
中文说明

本轮摘要 — PR #12016(issue #12006)

提交:10540b602b fix(ci): correct the /proc/locks witness and vet lock openability in the resolver (#12006)

第 11 轮的两条 Critical 发现均已复现并从根因修复,且各自补齐了发现所要求的测试覆盖。没有拒绝、延后或升级任何反馈。无合并冲突(--conflict false)。

[rc:4058048832] Critical R7-1 —— /proc/locks 见证错用 PID 字段 —— 已解决

已复现。 /proc/locks 的行格式为 <id>: FLOCK ADVISORY WRITE <pid> <major>:<minor>:<inode> <start> <end>:$5 是持有者 PID,$6 才是 设备号:inode 三元组(已对照活动 flock 的 /proc/self/fdinfo/8 核实)。已发布的 awk 切分的是 $5,因此 a[3] 恒为空,该见证永远无法报告「已被持有」——恰好在这个守卫为之存在的「不可读 root 属主锁」路径上失效(flock 探测以 66 退出)。对已发布脚本做端到端探测:在一个 0000 锁上放置真实活动持有者(flock 探测确实以 66 退出;由于本沙箱屏蔽了 /proc/locks,见证行通过 awk 垫片喂入),解析器判定共享目录健康、打印 ::warning::unlinked stale unwritable lock file …,并 unlink 了仍被活动持有的 inode(事后 stat:inode 已不存在)。

修复。 lock_held 现在用 stat -c %D/%i 构造 major:minor:inode(十六进制运算放在 bash 而非 awk——strtonum 仅 gawk 支持),并以 $6 整体比对,设备号与 inode 一起比较,避免仅按 inode 在其他文件系统上误认幻影持有者。按发现的约束,daemon 锁路径、解析器的命令行形态、以及 heal 只使用 >> 探测的方式均未改变。修复后的同一探测:作业私有回退、stderr 输出 is not writable and is locked by a live process、被持有 inode 保持原样。

测试。 按发现要求加在 scripts/tests/resolve-ci-lock-dir.test.js 的 refuses to unlink a lock a live process holds 旁边:

  • refuses to unlink an unreadable lock a live process holds——要求的 0000 用例(同样的 exec 8>> + 独占 flock 持有者,以 isRoot || !hasFlock || !hasProcLocksWitness 门控),断言 stdout 为作业私有回退、stderr 含持有者警告、inode 不变、且竞争者的 flock --nonblock --exclusive 仍失败。对所要求断言有一处刻意偏离:竞争者探测在把权限恢复为 0600 之后才执行,因为 0000 文件对任何打开都以 EACCES 拒绝,探测根本走不到它所要度量的 flock。门控中也有一处刻意新增:hasProcLocksWitness(持锁后自查的能力探测)取代了裸的 existsSync('/proc/locks'),因为沙箱可能把 /proc/locks 呈现为空的掩码文件——本运行器正是如此——此时见证环境根本无法存在,该用例会断言一个不可能产生的否决。在常规 Linux CI(以及该脚本所面向的 pool 运行器)上该能力探测通过、用例正常运行。
  • heals an unreadable lock no live process holds——要求的反向同胞用例(无人持有的 0000 锁仍必须被治愈),它正是抓住「恒判定为已持有」变体的用例。

变异探测。

  • 恒判定为已持有的见证变体(在 awk 之后植入 return 0):反向同胞用例转红(Tests 1 failed | 24 skipped),之后已逐字节还原。
  • $6→$5 回退:本沙箱的 /proc/locks 被掩码,已提交的正向用例在此按能力跳过,因此改用 awk 垫片驱动真实脚本做端到端回退验证:$5 变体 unlink 了被活动持有的 0000 锁并判定共享目录健康(inode 消失),而修复后的脚本否决并保持 inode——这正是已提交测试在 /proc/locks 正常的环境中所断言、并在回退时转红的性质。

[rc:4058048841] Critical R11-1 —— 调用方随后 open(),闸门却只测权限位 —— 已解决

已复现。 先写下发现要求的两个 prune 步骤用例并对改动前的代码运行:两者均红。daemon 锁名上放置 runner 属主的 0755 目录时,步骤打印 Docker cleanup skipped because the shared daemon is active(EISDIR 打开失败被 elif 条件在 bash -e 下吞掉,被谎报为 daemon 活动);放置目标位于同一可写目录内的悬空符号链接时,破坏性的 docker image prune --all --force 在一个不互斥任何东西的锁后面运行了。

修复(一次结构性收口,而非逐入口打补丁)。 在 resolve-ci-lock-dir.sh 的探测循环中、权限位分支之前:凡存在(或是悬空符号链接——-e 会跟随链接,所以用 -L 观察)但不是常规文件的名字一律拒绝;在既有 heal 之后:用 (exec 9>>…) 审查调用方紧接着要做的那次打开——追加模式,使探测永远不会截断活动进程持有 flock 的 inode——并以名字仍然存在为门,使得刚被 heal 掉的名字不会在调用方自己打开之前被重建(保持既有 existsSync(...) === false 的 heal 见证为绿)。常规文件审查正是这次打开有界的保证(FIFO 会把它挂住)。在 e2e.yml 的 prune 步骤中,exec 9> 从 elif 条件中移出,成为独立的 if ! exec … 分支并带显式的 cannot open … for writing 跳过行,bash -e 再也不会把打开失败折叠进 shared daemon is active 消息;解析器回退时的跳过原因现在点名 daemon 锁路径。scripts/tests/e2e-workflow.test.js 中逐字节锚定的 pin 已更新为新形态(if ! exec 9>… + elif flock --nonblock 9),断言数量不变(一个 pin 变为两个)。e2e.yml 仍在尺寸棘轮之内(35969 B,基线 33227 B + 4096 B 余量)。

测试。 加在 describe('prune step daemon lock discipline') 下,均驱动 runPruneStep(world, { withResolver: true }),且在改动前均为红:

  • skips the labelled prune when a directory squats on the daemon lock name(默认 0755,而非既有用例使用的 0500),以及
  • skips the labelled prune when a dangling symlink squats on the daemon lock name(目标位于同一可写目录内),

每个都断言输出不含 shared daemon is active、跳过行点名 docker-sandbox-daemon.lock、docker 日志中没有 image prune --all,并无条件的 dangling prune 仍运行;symlink 用例还锚定了占位者只被拒绝、绝不被修复性清除。

变异探测。 把解析器的审查回退为裸的 [[ -e && ! -w ]] 闸门会使 symlink 用例转红。目录用例在这一单独回退下保持绿色,因为步骤侧的打开分支——该发现要求的另一处修改——能独立地用真实的跳过行捕获 EISDIR;而对完全改动前的代码,两个用例均为红(已直接验证)。

给评审者的说明

  • 评审延后区中的六条 [probe] 项由评审记录为对本轮不阻断,本轮未处理。其中一条(仅按 inode 的幻影持有者缺口)已作为 R7-1 修复的承载部分一并关闭;其余(lock_held 的检查-再行动竞态、/proc/locks 分支覆盖、banner 的永久人工语义、分支侧的带标签 prune、严格前缀 pin)仍按记录保持延后。
  • 未触碰本 PR 自身足迹之外的 CI 设施;四个改动文件均已在 PR diff 中。没有删除、禁用或削弱任何测试——仅有一个 pin 字符串随 prune 步骤新形态更新,且断言数量增加。

验证

  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/resolve-ci-lock-dir.test.js —— 24 通过,1 跳过(0000 活动持有者用例在本沙箱被掩码的 /proc/locks 上按能力跳过;在常规 Linux CI 上运行)
  • npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/e2e-workflow.test.js —— 38 通过
  • npx vitest run … scripts/tests/workflow-size.test.js scripts/tests/no-ak-integration-ci.test.js(与上述两文件同一配置)—— 4 个文件,290 通过,1 跳过
  • 新 prune 步骤用例对改动前代码 —— 两者均按复现的缺陷失败(先红)
  • 变异探测 —— 恒已持有见证:反向同胞转红;解析器审查回退:symlink 用例转红;经 awk 垫片对真实活动持有者做 $5 回退:变体 unlink 了活动持有的 inode,修复版保持(细节见上)
  • bash -n .github/scripts/resolve-ci-lock-dir.sh —— 语法通过
  • 对全部四个改动文件运行 npx prettier --check —— 干净(测试文件先经 --write;格式化后重跑测试套件仍为绿)
  • npm run typecheck —— 通过
  • npm run build —— 通过;由于本运行器对任何单条命令有 120s 上限而整体构建超过该上限,按 scripts/build.js 的逐包顺序逐步执行(generate → browser-use → core → channels → audio-capture → node-repl → acp-bridge → sdk-typescript → web-shell → web-templates → cli → settings-schema → qwen-live → vscode-ide-companion → chrome-extension → external-context ×2),每步 rc=0;可选的 sandbox 镜像构建(BUILD_SANDBOX=1)未运行,且 generate-settings-schema 无 diff 产出
  • npm run lint(eslint . --ext .ts,.tsx + eslint integration-tests)—— 通过;在同一 120s 上限下分块运行(scripts + integration-tests、packages/cli、packages/core、其余全部 packages/integrations、根级文件),每块 rc=0

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。

🧵 Resolved all 2 selected review thread(s). · 已关闭全部选中的 2 条评审线程。

Re-review when you have a moment. After round 10 this bot stops and leaves the PR for a human. · 有空请复审;第 10 轮后本 bot 停止并将 PR 交给人工。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.24.2

@yiliang114

Copy link
Copy Markdown
Collaborator

Closing this — the linked issue #12006 was closed as not planned in favor of #11991, which prescribes a different remedy for the lock-dir problem. This diff also predates the pnpm migration (#11859) and no longer applies to the current e2e.yml. If the unwritable sandbox lock dir recurs on self-hosted runners, the #11991 direction is the one to implement.

@yiliang114 yiliang114 closed this Sep 21, 2026
auto-merge was automatically disabled September 21, 2026 06:36

Pull request was closed

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

review/self-reported The linked issue was opened by the PR author (self-reported)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Main CI failed: E2E Tests on aff26a8f36df

3 participants