Repository navigation
fix(ci): tolerate unwritable docker sandbox lock dir on self-hosted runners (#12006) - #12016
qwen-code-dev-bot wants to merge 14 commits into
Conversation
Autofix E2E Report — issue #12006IssueMain-branch E2E run 35069321648 on Diagnosis (evidence)
ReproductionDocker is not available in this autofix container, so the exact CI leg cannot run here. Instead the failure was reproduced at bash level by executing the real
FixBoth lock-opening call sites now resolve the lock directory through a new helper, The heal step itself is unchanged — it still restores full coordination where it has the privileges. Regression coverage
Verification
中文说明Autofix E2E 报告 — issue #12006问题main 分支在 诊断(证据)
复现本 autofix 容器中没有 docker,无法运行完全一致的 CI 分支。因此改为在 bash 层面复现:以真实
修复两个打开锁的调用点不再硬编码 修复步骤本身未改动——在有权限的环境中它仍能恢复完整协调。 回归覆盖新增的 验证
🧠 Handled by Qwen Code · model/模型 |
…#12006) Review round on #12016: - Stub flock(1) in the docker-leg test harness: it is util-linux and absent on macOS, where the new cases would fail the nightly lane. - The resolver now takes the lock names the caller will open and probes only those, so a foreign unwritable lock (sdk-java-tests.lock) no longer vetoes a usable shared dir. - The resolver header no longer claims the fallback is coordination-safe: RUNNER_TEMP is per-job, so coordination with sibling jobs, the root cleanup timer, and the release lane is lost; the fallback also lands on GITHUB_STEP_SUMMARY instead of only in a step log. - The always()-gated prune step tolerates a missing checkout (resolver absent) by falling back to the job-private dir inline. - Diagnostics name the offending path per probe failure instead of asserting an unverified root-owned cause. - Pin the fail-loud resolver branch and the exec 8> coordinator lock's directory, both previously unwitnessed. Co-authored-by: Qwen-Coder <[email protected]>
|
🤖 Addressed the latest review feedback (round 1/10). What changed, and what I pushed back on: · 已处理最新评审反馈(第 1/10 轮)。改动内容与我反驳保留之处如下: Autofix review round — PR #12016Addressed all 7 findings from the latest review (1 Critical + 6 Suggestions). No declines, no deferrals, no open questions. FindingsR1-1 [Critical] — docker-leg test cases fail where
|
|
🔀 Base updated: red check(s) [Lint & Static (ubuntu-latest, Node 22.x)] pass on current main — merged current main via update-branch; CI will re-run. 中文说明🔀 已更新 base:红色检查 [Lint & Static (ubuntu-latest, Node 22.x)] 在当前 main 上通过 —— 已通过 update-branch 合入当前 main,CI 将重新运行。 |
…12006) The prune step's 127 guard fell back to a job-private lock directory, where flock --nonblock always "wins" against a file no other process opens, so the labelled prune ran with its only cross-job exclusion defeated. Probe the shared dir inline when the resolver cannot run (it needs no checked-out file), and skip the labelled prune whenever the daemon lock would not land on the shared dir; the dangling prune stays unconditional because dangling images are untagged and unreferenced. The leg now resolves the daemon lock independently of the build-family locks, so a poisoned build lock can no longer move the daemon lock off the shared path the prune step, the host cleanup timer and the release lane all coordinate on. Also pins the resolver's "cannot create" diagnostic and witnesses that the test leg still fails loudly when no lock directory can be created.
|
🤖 Addressed the latest review feedback (round 2/10). What changed, and what I pushed back on: · 已处理最新评审反馈(第 2/10 轮)。改动内容与我反驳保留之处如下: Autofix round summary — PR #12016 (issue #12006)Commit: All seven open inline findings — two Criticals (R1-4 twice, R3-1) and two Suggestions (R1-8 twice, R2-1 twice) — are addressed in one coherent change: the daemon-lock coordination channel is now resolved and used identically by both call sites, and every new guard carries a mutation-proven witness. R1-4 (Critical, e2e.yml:335 — job-private fallback defeats the prune exclusion) — implementedThe Pins moved in the same change: R3-1 (Critical, e2e.yml:333 — leg and prune resolve the daemon lock with different name sets) — implementedThe leg now resolves the coordination channel on its own:
New witnesses: a resolver-level case ( R1-8 (Suggestion, resolve-ci-lock-dir.sh:36 —
|
) Address review feedback on the lock-dir fallback: - The prune step blanked a job-private resolver result but still logged "skipped because the shared daemon is active", and its resolver-absent inline probe checked only the directory, so a root-owned 0400 lock file failed `exec 9>` with EACCES inside the if-condition (no bash -e abort) and the step exited 0 with that false line as its only output. The step now probes the lock file inline, carries the cause it actually observed (naming the lock dir or file), and announces the unconditional dangling prune separately so "skipped" scopes to the labelled one. Its resolver call blanks GITHUB_STEP_SUMMARY: the step discards a job-private result and opens no lock in it, so the fallback banner would lie there (the leg's calls keep it). - The resolver's fallback warning and summary claimed "docker build/prune coordination is lost" even when only a build-family lock fell back; both now scope the loss to the locks the call actually probed ($*). - Tests: pin the unwritable-dir cause text at both reach points (and fix the comment that claimed it was pinned), add the success-direction witness driving the resolver and the step's shared-dir guard together so a one-token drift between them reddens, and assert the opened-locks == probed-locks invariant as a derived set so a lock added to an exec line without a probe-list entry fails the suite.
|
🤖 Addressed the latest review feedback (round 3/10). What changed, and what I pushed back on: · 已处理最新评审反馈(第 3/10 轮)。改动内容与我反驳保留之处如下: Addressed review feedback — PR #12016 roundAll five inline findings from the latest automated review were verified against the code, reproduced where they claimed wrong behavior, and resolved in commit Dispositions[rc:4031664803] R4-1 — prune step asserted causes it never verified — RESOLVEDVerified before implementing: on the resolver-absent path the inline probe tested only Changes in
Tests: retargeted [rc:4031664817] R4-2 — no success-direction witness for the resolver/step path agreement — RESOLVEDVerified: [rc:4031664822] R4-3 — fallback messages claimed prune coordination loss on build-only fallback — RESOLVEDVerified: the shared template printed [rc:4031664833] R4-4 — unwritable-dir cause text unpinned — RESOLVEDVerified: cause 2 ( [rc:4031664844] R4-5 — opened-locks == probed-locks invariant pinned only as literals — RESOLVEDVerified: a lock added to an Mutation probes (guard removed → matching case reddens → restored → green)
Notes
Verification
中文说明已处理评审反馈 —— PR #12016 本轮最新一轮自动评审的 5 条行内发现全部先对照代码核实、对声称行为错误的部分做了复现,并在提交 逐条处置[rc:4031664803] R4-1 —— 清理步骤断言了从未验证的成因 —— 已解决实现前已核实:在解析器缺失路径上,就地探测只对目录做
测试:改写了 [rc:4031664817] R4-2 —— 解析器与步骤路径约定缺少成功方向见证 —— 已解决已核实: [rc:4031664822] R4-3 —— 仅构建族回退也声称清理协调丢失 —— 已解决已核实:共享模板在只有构建族锁回退时也打印 [rc:4031664833] R4-4 —— 目录不可写成因的文本未被锚定 —— 已解决已核实:成因 2( [rc:4031664844] R4-5 —— 「打开的锁 == 探测的锁」不变量只有字面量锚定 —— 已解决已核实:给 变异探针(拆除守卫 → 对应用例变红 → 还原 → 复绿)
说明
验证
Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。 🧵 Resolved all 5 selected review thread(s). · 已关闭全部选中的 5 条评审线程。 Re-review when you have a moment. After round 10 this bot stops and leaves the PR for a human. · 有空请复审;第 10 轮后本 bot 停止并将 PR 交给人工。 🧠 Handled by Qwen Code · model/模型 |
|
⏸️ AutoFix paused by a review convergence signal: 3 consecutive automated review round(s) signalled machine-readable non-convergence (codes observed since the last maintainer response, or the window start if none: batch-fixes, root-cause-triage, stem-surface) — the reviewer keeps finding new defects at a rate that is not falling while the loop keeps widening the diff, so another automatic round is unlikely to converge this PR. The loop resumes once a maintainer responds on this PR (a review or comment counts, and steers the next rounds), and pauses again if the signal persists for 3 more round(s). Alternatives: split the recurring cluster into its own PR, batch the remaining fixes into one push, comment 中文说明⏸️ AutoFix 已因评审收敛信号暂停:3 轮自动评审连续发出机器可读的不收敛信号(自上次维护者响应以来观察到的信号码;若无响应则自窗口开始:batch-fixes, root-cause-triage, stem-surface)——评审仍在以不降的速率发现新缺陷,而循环在继续扩大 diff,再跑一轮自动修复难以收敛本 PR。维护者在本 PR 上作出回应后循环自动恢复(评论或评审均可,并将作为后续轮次的指引);若信号再持续 3 轮会再次暂停。可选做法:把反复出问题的簇拆成独立 PR、把剩余修复攒成一批一次推送、评论 |
|
@qwen-code /triage |
Maintainer verification — built the real environment locally and drove this change through itVerified at head How it was verifiedNot by re-reading the diff. A real Linux host was reconstructed and the three step bodies were executed on it, byte-for-byte as the workflow ships them:
1. The failure reproduces on main, and the PR fixes itThe main arm emits the production log of run 35069321648 line for line, including the line number: Two details of the design hold up under test, and neither is pinned by reading alone:
2. The incident is wider than the linked issuePulling every From 2026-09-17 04:11Z onward the same hosts run the leg green, so the poisoned state ended within about 21 hours. Whether a human cleared it or something else did is not visible from the logs — worth noting because the resolver header asserts "The state does not heal itself: clearing it needs a human on the host," and production neither confirms nor refutes that. 3. On a healthy host the change is not quite byte-for-byteWith no sibling leg running, every recorded stream is identical between the arms except one added stdout line — same docker argv, same locks, same exit codes. But with a sibling leg live, the arms diverge:
The PR moves 4. The standing Critical (R7-1) — confirmed, and it is not merely theoreticalDriving the real
Against real Docker: an image the leg has just accepted via Where R7-1's premise is only half-supported. It rests on the poison being a root-owned lock file inside a runner-owned directory, which the runner can
And the two shapes are indistinguishable in production: the heal warning, the leg's R7-1's proposed repair was implemented and measured (narrow
It works, it does not weaken the unrepairable case, and it costs 3 existing cases that currently assert "unwritable lock file → fallback" and would need re-pointing at the unrepairable shape. That is a real follow-up, not a reason to hold this PR: on the shapes it cannot repair, the fallback is still the only thing standing between the host and a dead leg. 5. The other open findings
Mutation coverage of what the PR actually protects is otherwise strong — 7 of 8 mutations killed, including every resolver-core one (3–9 cases each). 6. Review gaps this closesThe last review disclosed four gaps. Three are now closed:
RecommendationMerge. The change fixes a failure that is reproducible, recently recurrent across two pool hosts, and fatal to the leg before a single test runs; on a healthy host it costs one log line and one prune that the host's own timer already performs unlocked. The standing Critical is accurate about the cost on a poisoned host, but the alternative on such a host is a leg that cannot run at all. Two follow-ups worth filing rather than blocking on:
Two corrections to make in the description before merge: the "byte-for-byte" sentence (§3, both languages), and 中文说明维护者验证 —— 在本地搭建真实环境并把本改动放进去跑验证对象为 head 验证方式不是靠读 diff,而是重建了一台真实的 Linux 主机,并按工作流原样逐字执行三个步骤体:
1. 故障在 main 上复现,PR 修好了它main 臂逐行重现了 run 35069321648 的生产日志,连行号都一致: 设计中的两个细节经测试成立,且都不是读代码能钉住的:
2. 事故范围比关联 issue 更广把 2026-09-16 至 09-20 之间所有 从 09-17 04:11Z 起,同样的主机上该分支恢复绿色,也就是说中毒状态在约 21 小时内结束了。是人工清理还是别的原因,日志看不出来——值得一提,因为解析器文件头断言「该状态不会自行愈合:清除它需要有人登录主机」,而生产数据既不能证实也不能证伪这一点。 3. 健康主机上并非完全逐字节一致没有兄弟分支在跑时,两臂记录到的所有流只差一行新增的 stdout——docker argv、锁位置、退出码全部相同。但当有兄弟分支正在运行时,两臂分叉:labelled prune 都跳过;dangling prune 在 main 上跳过,在 PR 上执行。 PR 把 4. 未决 Critical(R7-1)——成立,而且不只是理论以 root 身份对着一个活着的 leg 运行真实的
对真实 Docker:一个刚被 leg 通过 R7-1 前提只有一半被支持。 它成立的前提是毒物为 runner 属主目录里的 root 属主锁文件,这种情况 runner 无需 sudo 即可 而这两种形态在生产中无法区分:heal 的 warning、leg 的 R7-1 提议的修复已实现并实测(只 unlink 探测判定为不可写的锁名,重测,失败则回退):root 属主锁文件的世界里 leg 绿且 daemon 锁留在共享目录、毒物被清成 runner 属主 0644;root 属主目录的世界里正确回退;健康世界不变。它有效、不削弱不可修复的情形,代价是现有 3 个用例(当前断言「锁文件不可写 → 回退」)需要改指到不可修复的形态。那是一个真实的后续项,而不是拦住本 PR 的理由:在它修不了的形态上,回退仍是主机与「分支直接死掉」之间唯一的屏障。 5. 其余未决发现
除此之外,对 PR 真正保护的东西,变异覆盖是扎实的——8 个变异杀死 7 个,解析器核心的每一个变异都被杀(各 3–9 个用例)。 6. 本次填补的审查缺口上一轮评审披露了四处缺口,现已填补三处:
建议合并。 该改动修复的是一个可复现、近期在两台池主机上复发、且在任何测试运行前就让分支死亡的故障;在健康主机上的代价是一行日志和一次主机定时器本来就不持锁执行的 prune。未决 Critical 对中毒主机上的代价描述准确,但在这样的主机上,另一个选项是分支根本跑不起来。 两个值得建 issue 而非拦住本 PR 的后续项:
合并前建议更正描述中的两处:「逐字节一致」那句(§3,中英文各一处),以及 |
|
@qwen-code /triage |
…ck (#12006) The resolver vetoed the whole shared dir when a caller-named lock file was unwritable, but unlink permission lives on the containing directory: a root-owned lock file inside a runner-writable dir is exactly the #12006 shape, and the runner can clear it without sudo. Unlink only the files the probe found unwritable, re-test, and fall back only when the repair fails, so a repairable host keeps cross-job coordination (sibling legs, the root qwen-docker-cleanup timer, the release-lane build mutex) instead of losing it until a human logs in. Co-authored-by: Qwen-Coder <[email protected]>
|
🤖 Addressed the latest review feedback (round 4/10). What changed, and what I pushed back on: · 已处理最新评审反馈(第 4/10 轮)。改动内容与我反驳保留之处如下: Autofix round — PR #12016 (issue #12006)Growth audit (required this round)Verdict: sound — recorded in
Feedback dispositionsR7-1 (Critical) — rc:4046529395, rc:4049746075, rc:4054490664; standing in reviews rv:5247538768, rv:5251440468, rv:5257402698 — implementedClaim: the unconditional veto abandons the shared lock dir when a named lock file is unwritable, but in the #12006 shape (root-owned lock file inside a runner-writable dir) the runner can Reproduced on the pre-round code before implementing (probe, this runner, non-root): shared dir writable + Fix ( Tests (
Mutation probes (both run, both killed):
Maintainer verification comment (ic:5747435862, @wenshao)
Deferred non-Critical sectionNot work — audit record under critical-only mode; no code, thread, or reply actions taken on those items. Verification
Notes
中文说明Autofix 本轮结果 — PR #12016(issue #12006)增长审计(本轮必需)结论:sound——已在任何编辑之前写入
反馈处理R7-1(Critical)—— rc:4046529395、rc:4049746075、rc:4054490664;并立于评审 rv:5247538768、rv:5251440468、rv:5257402698 —— 已实现主张:当一个被点名的锁文件不可写时,无条件否决放弃了整个共享锁目录;但在 #12006 的形态下(runner 可写目录里的 root 属主锁文件),runner 无需 sudo 即可 实现前已在改动前代码上复现(本机非 root 探测):共享目录可写 + 修复( 测试(
变异探测(两个均已执行、均被杀死):
Maintainer 验证评论(ic:5747435862,@wenshao)
延后的非 Critical 区段非工作项——仅处理 Critical 模式下的审计记录;未对这些条目做代码、线程或回复动作。 验证
说明
Deferred non-Critical feedbackCritical-only mode is active: the PR's diff grew src 33 / test 410 net lines beyond this counting window's baseline (budgets: 400/400). The workflow excluded the non-Critical feedback below from this round's actionable sections; the items remain open for human follow-up. Maintainer feedback continues to flow unaffected during a growth-only engagement (the per-author batch budget applies only after 5 change-producing rounds). (
中文说明已进入仅处理 Critical 的模式:本计数窗口内 diff 净增长已达 源码 33 / 测试 410 行(预算 400/400)。本轮可执行区域已排除下方非 Critical 反馈;这些条目保持开放,留待人工跟进。纯增长触发期间维护者反馈照常流动(按作者的批次预算仅在完成 5 个产生改动的轮次后生效)。(评论 Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。 🧵 Resolved all 1 selected review thread(s). · 已关闭全部选中的 1 条评审线程。 Re-review when you have a moment. After round 10 this bot stops and leaves the PR for a human. · 有空请复审;第 10 轮后本 bot 停止并将 PR 交给人工。 🧠 Handled by Qwen Code · model/模型 |
|
🔀 Base updated: red check(s) [Test (ubuntu-latest, Node 22.x), Lint & Static (ubuntu-latest, Node 22.x), Integration Tests (no-AK, No Sandbox), web-shell E2E Smoke (ubuntu-latest, Node 22.x)] pass on current main — merged current main via update-branch; CI will re-run. 中文说明🔀 已更新 base:红色检查 [Test (ubuntu-latest, Node 22.x), Lint & Static (ubuntu-latest, Node 22.x), Integration Tests (no-AK, No Sandbox), web-shell E2E Smoke (ubuntu-latest, Node 22.x)] 在当前 main 上通过 —— 已通过 update-branch 合入当前 main,CI 将重新运行。 |
|
🤖 AutoFix updated a stale base — the fix did not pass verification, but this PR was behind What I found before stopping: See the Qwen Autofix agent step logs for model/tool output. 中文说明🤖 AutoFix 更新了一个过期的 base —— 修复未通过验证,但本 PR 落后于 Run log: https://github.com/QwenLM/qwen-code/actions/runs/35497432176 🧠 Handled by Qwen Code · model/模型 |
|
🤖 Reviewed the latest feedback — no changes needed. Why, point by point: · 已审阅最新反馈——无需改动。逐点说明原因如下: Autofix round — PR #12016 (issue #12006): no code change this roundWhy no changeEvery actionable finding is already resolved in code at the current head (
Resolved this round (re-verified, fixed by an earlier commit)
Deferred to the follow-up queue (recorded, threads left open with a reply)Per the maintainer's recommendation and the convergence posture, no Suggestion-level item is implemented in this PR:
All recorded in Needs the maintainer (this bot has no GitHub writes)Two PR-description corrections from the verification comment remain open — they require editing the PR body, which only a human or the PR workflow can do:
Note on the two triage-verify failuresIssue comments 5747219213 and 5747495869 report the sandboxed-verification harness failing at Verification
No commit was created; the working tree is unchanged. 中文说明Autofix 本轮结论 —— PR #12016(issue #12006):本轮无代码改动为什么不做改动在当前 head(
本轮解决(重新验证,由先前提交修复)
转入后续跟进队列(已记录,线程回复后保持未解决)按维护者建议与收敛姿态,本 PR 不再实现任何建议级条目:
以上全部记录于 需要维护者处理(本 bot 无 GitHub 写权限)验证评论中要求的两处 PR 描述更正仍未完成——它们需要编辑 PR 正文,只能由人工或 PR 工作流执行:
关于两次 triage-verify 失败的说明issue 评论 5747219213 与 5747495869 报告沙箱验证装置在任何验证开始前的 验证
本轮未创建提交;工作树未改动。 Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。 🧵 Resolved all 1 selected review thread(s). · 已关闭全部选中的 1 条评审线程。 🧠 Handled by Qwen Code · model/模型 |
|
🕐 Review received — an automatic review of the current head is still running, so this round is held until it lands (a push now would cancel it and discard its work, #8888). Your feedback stays queued for the next eligible round. 中文说明🕐 已收到评审 —— 当前 head 上仍有一轮自动 review 在运行,本轮暂缓(现在推送会取消该 review 并丢弃其工作,#8888)。反馈保持排队,等待下一次可运行的轮次处理。 |
Co-authored-by: Qwen-Coder <[email protected]>
|
🤖 Addressed the latest review feedback (round 2/10). What changed, and what I pushed back on: · 已处理最新评审反馈(第 2/10 轮)。改动内容与我反驳保留之处如下: Autofix round — PR #12016 (issue #12006)Commit: Feedback addressed[rc:4057016408] — Critical R7-1: the in-place heal unlinks a lock file on the sole evidence of a permission bit — implementedReproduced first (source-blind). With a live holder mirroring the root Fix (
Also per the finding: the successful heal now emits a Tests (
Mutation probes (guard deleted → test red → restored → green):
Both changed tests therefore fail on the pre-round code and pass on the round's code. Not addressed (out of scope this round)
Verification
中文说明Autofix 本轮报告 — PR #12016(issue #12006)提交: 已处理的反馈[rc:4057016408] — Critical R7-1:就地治愈仅凭权限位就 unlink 锁文件 —— 已实现先复现(source-blind)。 用一个完全复刻 root 修复(
同时按发现要求:治愈成功后在 stderr 发出点名被 unlink 文件的 测试(
变异探针(删除守卫 → 用例转红 → 恢复 → 转绿):
因此两个改动过的用例在本轮之前的代码上都会失败、在本轮代码上通过。 本轮未处理(超出范围)
验证
Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。 🧵 Resolved all 1 selected review thread(s). · 已关闭全部选中的 1 条评审线程。 Re-review when you have a moment. After round 10 this bot stops and leaves the PR for a human. · 有空请复审;第 10 轮后本 bot 停止并将 PR 交给人工。 🧠 Handled by Qwen Code · model/模型 |
…the resolver (#12006) Co-authored-by: Qwen-Coder <[email protected]>
|
🤖 Addressed the latest review feedback (round 3/10). What changed, and what I pushed back on: · 已处理最新评审反馈(第 3/10 轮)。改动内容与我反驳保留之处如下: Round summary — PR #12016 (issue #12006)Commit: Both Critical findings from round 11 were reproduced and fixed at the root, each with the test coverage the finding asked for. No feedback was declined, deferred, or escalated. No merge conflicts ( [rc:4058048832] Critical R7-1 —
|
|
Closing this — the linked issue #12006 was closed as not planned in favor of #11991, which prescribes a different remedy for the lock-dir problem. This diff also predates the pnpm migration (#11859) and no longer applies to the current e2e.yml. If the unwritable sandbox lock dir recurs on self-hosted runners, the #11991 direction is the one to implement. |
Pull request was closed




What this PR does
The self-hosted Linux docker E2E leg opens its host-side flock files (daemon prune-exclusion, per-commit build coordinator, build mutex) under
${HOME}/.cache/qwen-code-ci. When a root-owned leftover makes that directory unusable, everyexec 9>-style open dies with EACCES before a single test runs. This PR routes both lock-opening call sites — the E2E runner script and the workflow's image-prune step — through a small resolver script that prints the shared directory when it is genuinely usable (directory writable and every existing lock file writable) and otherwise prints a job-private fallback under the job's temp dir, with a::warning::in the log. On healthy hosts the resolved path is the same shared directory as before, so the locking protocol is unchanged; on a degraded host the leg now runs with per-job coordination instead of dying in under a second.Why it's needed
Main-branch E2E run 35069321648 (tracked in #12006) failed the
sandbox:dockerleg'sRun E2E testsandPrune dangling docker imagessteps in under one second each. The job'sRestore workspace ownershipstep had already warnedcould not restore CI cache ownership: itschown … || sudo -n chown …chain cannot repair a root-owned lock directory on a runner whose user has no passwordless sudo (a non-root process can never chown another user's files). This is the same failure class as #11990 recurring on a host where the heal is powerless, so the workflow must tolerate the unwritable directory instead of depending on healing it.Losing cross-job coordination on such a host is safe: the sandbox image build is idempotent (a duplicate concurrent build wastes work but converges on the same tag), and both prune paths are filtered by label and
until=24h, so they cannot remove an image a running leg is using — the host cleanup service already relies on the age gate rather than the lock for exactly this reason.Reviewer Test Plan
How to verify
The change is CI-runner shell behavior; the review surface is the new resolver and its two call sites.
.github/scripts/resolve-ci-lock-dir.shprints the shared dir only when it is writable along with every existing*.lock, otherwise falls back to${RUNNER_TEMP}/qwen-code-ci-locksand warns on stderr (stdout stays a clean single path for$(...)capture).npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/resolve-ci-lock-dir.test.js scripts/tests/e2e-workflow.test.js— the new suite executes the real runner script's docker branch against an unwritable lock file (the exact Main CI failed: E2E Tests on aff26a8f36df #12006 reproduction: exit 1 EACCES pre-fix, green post-fix) and against a healthy home; the workflow suite pins both call sites to the resolver.::warning::about the job-private fallback and the docker leg runs normally; a healthy host is byte-for-byte the old behavior.Evidence (Before & After)
N/A — CI infrastructure only. Before:
run-e2e-tests.sh: line 46: …/docker-sandbox-daemon.lock: Permission denied, step exit 1 in <1s (reproduced locally with achmod 0400lock file). After: the same poisoned setup completes the leg on a job-private lock dir with a warning.Tested on
Environment (optional)
Autofix container (no docker daemon): bash-level execution of the real runner script with stubbed
docker/npx, plus the repository's workflow-contract vitest suites. The real docker leg cannot run here; the next main-branch E2E run is the end-to-end confirmation.Risk & Scope
sandbox:dockerE2E leg (no docker daemon in the authoring environment); yamllint (no pip in the container — the two-line YAML edit preserves indentation and actionlint parses the file cleanly)..github/scripts/run-release-docker-integration.shandsdk-java.ymluse the same shared lock dir and keep the same exposure; fixing them is a deliberate follow-up to keep this PR scoped to the failing workflow. A pre-existingverify-capture.test.jsANSI-rendering failure exists in this container and reproduces identically on cleanHEAD; it is unrelated.Restore workspace ownershipheal step is kept as-is — where sudo exists it still restores full shared-dir coordination.Linked Issues
Fixes #12006
中文说明
本次 PR 的内容
自托管 Linux docker E2E 分支在
${HOME}/.cache/qwen-code-ci下打开宿主机 flock 文件(daemon prune 互斥、按提交构建协调、构建互斥)。当 root 属主的残留导致该目录不可用时,每个exec 9>式打开都会以 EACCES 失败,连一个测试都来不及运行。本 PR 将两个打开锁的调用点——E2E 运行脚本和工作流中的镜像清理步骤——改为通过一个小型解析脚本获取锁目录:共享目录确实可用(目录可写且所有现存锁文件可写)时输出它,否则输出作业临时目录下的作业私有回退目录,并在日志中打印::warning::。主机健康时解析结果与原先完全相同,锁定协议不变;主机受损时该分支以作业内协调运行,而不是在一秒内死亡。为什么需要
main 分支 E2E 运行 35069321648(由 #12006 跟踪)中,
sandbox:docker分支的Run E2E tests和Prune dangling docker images两个步骤都在不到一秒内失败。该作业的Restore workspace ownership步骤已经告警could not restore CI cache ownership:其chown … || sudo -n chown …链在 runner 用户没有免密 sudo 时无法修复 root 属主的锁目录(非 root 进程永远无法 chown 他人的文件)。这与 #11990 是同一故障类别在修复手段失效的主机上的复发,因此工作流必须容忍目录不可写,而不是依赖修复它。在这类主机上失去跨作业协调是安全的:sandbox 镜像构建是幂等的(并发的重复构建只是浪费算力,最终收敛到同一个 tag),而两条 prune 路径都带标签和
until=24h过滤,不可能删除正在运行的分支在用的镜像——宿主机清理服务本来就依赖时限而非锁来保证这一点。审查者测试计划
如何验证
本次改动是 CI runner 的 shell 行为;审查重点是新的解析脚本及其两个调用点。
.github/scripts/resolve-ci-lock-dir.sh仅在共享目录可写且所有现存*.lock可写时输出它,否则回退到${RUNNER_TEMP}/qwen-code-ci-locks并向 stderr 告警(stdout 保持为干净的单一路径供$(...)捕获)。npx vitest run --config ./scripts/tests/vitest.config.ts scripts/tests/resolve-ci-lock-dir.test.js scripts/tests/e2e-workflow.test.js——新套件用不可写锁文件真实执行 runner 脚本的 docker 分支(即 Main CI failed: E2E Tests on aff26a8f36df #12006 的精确复现:修复前退出码 1 EACCES,修复后转绿)以及健康 home 的情形;工作流套件锚定两个调用点都经过该解析器。::warning::并正常跑完 docker 分支;健康主机与旧行为逐字节一致。前后对比证据
N/A——纯属 CI 基础设施。修复前:
run-e2e-tests.sh: line 46: …/docker-sandbox-daemon.lock: Permission denied,步骤 <1 秒退出码 1(已用chmod 0400锁文件在本地复现)。修复后:同样的中毒设置在作业私有锁目录上完整跑完该分支并输出一条警告。已测试平台
环境(可选)
autofix 容器(无 docker 守护进程):在 stub 的
docker/npx下真实执行 runner 脚本的 bash 级验证,外加仓库的工作流契约 vitest 套件。真实的 docker 分支无法在此运行;下一次 main 分支 E2E 运行是端到端确认。风险与范围
sandbox:dockerE2E 分支(编写环境无 docker 守护进程);yamllint(容器内无 pip——两行 YAML 改动保持缩进一致,actionlint 已干净解析)。.github/scripts/run-release-docker-integration.sh与sdk-java.yml使用同一共享锁目录、保留同样的暴露面;修复它们是有意留下的后续项,以保持本 PR 只覆盖出故障的工作流。容器内还存在一个既有的verify-capture.test.jsANSI 渲染失败,在干净HEAD上可同样复现,与本次改动无关。Restore workspace ownership修复步骤保持原样——在有 sudo 的环境中它仍能恢复完整的共享目录协调。关联 Issue
Fixes #12006