Repository navigation
fix(release): reclaim docker disk and gate the data root before the sandbox image build (#13479) #13481
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
base: main
Are you sure you want to change the base?
fix(release): reclaim docker disk and gate the data root before the sandbox image build (#13479) #13481
Changes from 5 commits
bf876c4
6aebce3
32e8b5a
8adfb17
2560f56
b6fae43
175d41d
5b6a321
46c54cf
fae808c
2bdbe7f
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -24,6 +24,31 @@ trap 'exit 1' INT TERM | |
| sandbox_revision="$(git rev-parse HEAD)" | ||
| sandbox_image="$(node -p "require('./packages/cli/package.json').config.sandboxImageUri")-release-${sandbox_revision}" | ||
|
|
||
| # The job-start disk floor gate predates this build: run 37374675168 passed | ||
| # it and the runner still died on ENOSPC 24 minutes into this step (#13479). | ||
| # Gate the build's own filesystem — the docker data root — at a build-sized | ||
| # floor, so a saturated host fails fast with a legible error and a re-run | ||
| # lands on an instance with headroom instead of the runner worker crashing | ||
| # mid-build. 8 GiB covers a cold builder stage (monorepo install + bundle | ||
| # layers) plus the final image with margin. Self-hosted only, like every | ||
| # other check-disk-floor.sh call site: an ephemeral hosted runner starts | ||
| # with an order of magnitude more free disk than this floor. | ||
| check_docker_data_root_floor() { | ||
| if [ "$RUNNER_ENVIRONMENT" != 'self-hosted' ]; then | ||
| return 0 | ||
| fi | ||
| if [ ! -f .github/scripts/check-disk-floor.sh ]; then | ||
| echo "::warning::docker data root floor gate skipped: .github/scripts/check-disk-floor.sh not present at this ref on ${RUNNER_NAME:-this runner}" | ||
| return 0 | ||
| fi | ||
| docker_root="$(timeout 60 docker info --format '{{.DockerRootDir}}' 2>/dev/null || true)" | ||
| if [ -z "$docker_root" ] || [ ! -d "$docker_root" ]; then | ||
| echo "::warning::docker data root floor gate skipped: docker data root '${docker_root:-<unreadable>}' is not a readable directory on ${RUNNER_NAME:-this runner}" | ||
| return 0 | ||
| fi | ||
| DISK_FLOOR_MIN_FREE_KB="${DISK_FLOOR_MIN_FREE_KB:-8388608}" bash .github/scripts/check-disk-floor.sh "$docker_root" | ||
| } | ||
|
|
||
| if [ "$RUNNER_ENVIRONMENT" = 'self-hosted' ]; then | ||
| mkdir -p "${HOME}/.cache/qwen-code-ci" | ||
| # Same protocol as e2e.yml: the host daemon lock is held shared for the whole | ||
|
|
@@ -51,7 +76,21 @@ if ! docker image inspect "$sandbox_image" > /dev/null 2>&1; then | |
| exit 1 | ||
| fi | ||
| fi | ||
| docker image prune --all --force --filter 'label=org.qwen-code.ci.sandbox=true' --filter 'until=24h' || echo "::warning::old CI sandbox image cleanup failed on ${RUNNER_NAME:-this runner}" | ||
| timeout 20m docker image prune --all --force --filter 'label=org.qwen-code.ci.sandbox=true' --filter 'until=24h' || echo "::warning::old CI sandbox image cleanup failed on ${RUNNER_NAME:-this runner}" | ||
|
Collaborator
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. [Suggestion] R4-1: The reclaim sequence added in this block is the pool's fourth copy of the same docker reclaim policy, and the four copies now disagree on which commands they run, on whether they are time-bounded, and on The contention premise is not speculative. Witness: The in-scope minimum is to bound the peer lane's twin — apply the same A shared helper must keep If the minimal variant lands, an assertion in 中文说明本块新增的回收序列是这个 runner 池上同一套 docker 回收策略的第四份副本,而四份副本在“执行哪些命令”“是否有时间上限”“是否带 争用前提并非猜测: 证据: 最小改动范围的做法是给对端那份也加上限:对 共享辅助脚本必须把 如果采用最小方案,需要补的测试在 — qwen3.8-max via Qwen Code /review (v0.25.0)
Collaborator
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Verified before deciding: Deferred to the follow-up queue rather than implemented here: bounding the E2E lane means editing the shared CI script every PR's E2E shards run on, which is CI machinery outside this release-lane PR's footprint — the same call your fuller-fix variant makes. The finding, the minimal fix, and the suggested What did land here, as your fallback option: this script now records the asymmetry next to the mutex — every daemon call it makes is time-bounded so a wedged GC cannot starve a peer lane's 30-minute wait, and the E2E lane's unbounded copy of the same prune under the same lock is noted as a cross-lane change tracked separately. 中文说明决定前已验证: 本次不直接实现,而是转入后续跟进队列:给 E2E 通道加上限意味着改动每个 PR 的 E2E 分片都在运行的共享 CI 脚本,属于本发布通道 PR 范围之外的 CI 设施——与您给出的完整修复方案(共享回收辅助脚本)同属跨通道改动。该发现、最小修复方案以及建议补在 按您给出的兜底方案,本脚本已在互斥锁旁记录了这种不对称:本脚本的每次守护进程调用都带时间上限,卡死的 GC 不会饿死对端通道 30 分钟的锁等待;E2E 通道在同一锁下的同款裁剪仍无上限,注释中已注明这是另行跟踪的跨通道改动。 |
||
| # The labelled prune cannot reach untagged images, and this lane passes | ||
|
qwen-code-dev-bot marked this conversation as resolved.
Outdated
|
||
| # --no-prune to the build: an image that went dangling after the daily | ||
| # 02:30 UTC sweep would otherwise never be reclaimed. | ||
| timeout 20m docker image prune --force --filter 'until=24h' || echo "::warning::dangling image cleanup failed on ${RUNNER_NAME:-this runner}" | ||
|
qwen-code-dev-bot marked this conversation as resolved.
Outdated
|
||
| # Image pruning does not reclaim BuildKit's intermediate install/build | ||
| # layers. The daily host sweep (ecs-runner/qwen-docker-cleanup) bounds them | ||
| # at 30 GB, but this lane builds at the end of the pool's day, hours after | ||
| # that sweep. Every daemon call in this branch is bounded so a slow daemon | ||
| # GC cannot hold the host build mutex past the E2E lane's 30-minute lock | ||
| # wait. No --keep-storage here: it is a reserve, not a quota, so on a host | ||
| # with less cache than the reserve — exactly the hosts this line exists | ||
| # for — it would reclaim nothing. | ||
| timeout 20m docker builder prune --all --force --filter 'until=24h' || echo "::warning::docker build cache cleanup failed on ${RUNNER_NAME:-this runner}" | ||
| check_docker_data_root_floor | ||
| # See e2e.yml: closing the lock descriptors in the child keeps a descendant | ||
| # that outlives this job from holding the lock. | ||
| npm run build:sandbox -- -s --no-prune -i "$sandbox_image" 7>&- 8>&- 9>&- | ||
|
|
@@ -67,6 +106,14 @@ if [ "$RUNNER_ENVIRONMENT" = 'self-hosted' ]; then | |
| exec 8>&- | ||
| fi | ||
|
|
||
| # Run 37374675168 actually died in the vitest phase, ~15 minutes after the | ||
|
Collaborator
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. [Suggestion] R4-3: Two point samples now cover three heavy phases, and the phase the incident actually died in is the one nothing measures. This gate is taken once, before two back-to-back vitest docker phases, and A host that clears this sample with just over 8 GiB free can cross zero during the cli phase — container writable layers, transcripts and the runner's own diagnostic log all land on that filesystem — and the worker dies on ENOSPC mid-phase with no step conclusion, so even the Witness: Attach the sampler #10394 established to this lane: start a 10-second Each sample must go to the job log as well as to the file — A new case in 中文说明现在用两个时间点采样去覆盖三个重负载阶段,而事故真正丧生的那个阶段恰恰没有任何测量。这个门只取一次样,位于两个连续执行的 vitest docker 阶段之前;而 一台以略高于 8 GiB 的余量通过本次采样的主机,可能在 cli 阶段期间归零——容器可写层、测试记录以及 runner 自身的诊断日志都写在同一个文件系统上——于是 worker 在阶段中途因 ENOSPC 死亡,步骤没有任何结论,连 证据: 请把 #10394 确立的采样器接到本通道上:在第一个 每个采样必须同时写入作业日志和文件—— 需要补的测试在 — qwen3.8-max via Qwen Code /review (v0.25.0)
Collaborator
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Verified before deciding: Deferred to the follow-up queue rather than implemented here: attaching the #10394 sampler means adding steps to a workflow this PR never touches — CI machinery outside this PR's footprint. The full design you specified (10-second loop over the docker root and The secondary ask — a third gate sample between the two vitest invocations — was considered and not added: a point sample there does not diagnose the incident shape (only the streamed sampler survives a worker crash), and each added gate is flake surface on the lane. It rides along in the same deferral record. 中文说明决定前已验证: 本次不直接实现,而是转入后续跟进队列:接入 #10394 确立的采样器需要向本 PR 从未改动过的工作流添加步骤,属于本 PR 范围之外的 CI 设施。您指定的完整设计(对 docker 根目录与 次要建议——在两个 vitest 调用之间再加一次门采样——经考虑后未采纳:在那里做一次点采样无法诊断事故形态(worker 崩溃时只有流式采样能留存),而且每增加一道门都会给通道带来抖动面。该点已一并记入同一条延期记录。
Collaborator
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. [Suggestion] R4-3: Still stands. Two point samples now cover three heavy phases, and the phase the incident actually died in is the one nothing measures. This gate is taken once, before two back-to-back vitest docker phases, and A host that clears this sample with just over 2 GiB free can cross zero during the cli phase — container writable layers, transcripts and the runner's own diagnostic log all land on that filesystem — and the worker dies on ENOSPC mid-phase with no step conclusion, so even the Witness (measured in the review worktree at head You verified this gap in comment 4211178534 and deferred it to the follow-up queue because the remedy adds steps to a workflow this PR never touches, which is a reasonable scoping call and is recorded in the deferred-findings issue. Nothing landed this round, so the entry carries forward rather than closing — the deferral record is now the only thing carrying it. The design you specified there is the fix: a 10-second Each sample must go to the job log as well as to the file, which is what 中文说明[建议] R4-3:仍然存在。现在是两次点采样覆盖三个重载阶段,而事故真正死掉的那个阶段没有任何东西在测量。这道门只取一次样,位于两个连续的 vitest docker 阶段之前,而 一台以刚好超过 2 GiB 的剩余空间通过本次采样的主机,可能在 cli 阶段降到零——容器可写层、日志输出以及 runner 自身的诊断日志都写在那个文件系统上——于是 worker 在阶段中途因 ENOSPC 死掉,步骤没有任何结论,连 证据(在 head 你已在评论 4211178534 中验证了该缺口并转入后续跟进队列,理由是补救措施需要向本 PR 从未改动的工作流添加步骤——这是合理的范围判断,也已记入延期发现 issue。本轮没有任何改动落地,因此该条目向前携带而不是关闭——现在延期记录是唯一携带它的东西。你在那里指定的设计就是修复方案:在第一次 每次采样必须同时写入作业日志与文件,也就是 — qwen3.8-max via Qwen Code /review (v0.25.0)
Collaborator
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Verified and still deferred to the follow-up queue: the remedy (a 10-second df sampler over the docker data root and RUNNER_TEMP started before the first npx vitest, stopped from the existing EXIT trap, each sample echoed to the job log as well as the file — the ci.yml:780 pattern, because the incident's worker crash loses a file-only artifact — plus an if: failure() upload-artifact step) adds steps to 中文说明已核实,仍延后至跟进队列:补救方案(在第一次 npx vitest 之前启动一个针对 docker 数据根目录与 RUNNER_TEMP 的 10 秒 df 采样循环,由现有 EXIT trap 停止,每条采样同时写入作业日志与文件——即 ci.yml:780 的做法,因为事故中 worker 崩溃会丢失仅存文件的产物——外加一个 if: failure() 的 upload-artifact 步骤)需要向 |
||
| # build finished: re-gate now that the build's peak and the image it leaves | ||
| # behind have landed on this filesystem. The gate sits outside the | ||
| # image-missing branch so a re-run that finds the image already cached on a | ||
|
qwen-code-dev-bot marked this conversation as resolved.
|
||
| # still-saturated host is gated before vitest writes container layers to the | ||
| # same filesystem. | ||
| check_docker_data_root_floor | ||
|
qwen-code-dev-bot marked this conversation as resolved.
Outdated
|
||
|
|
||
| # The package.json docker test scripts each rebuild the sandbox image. Run | ||
| # vitest directly here so this job reuses the image built above. | ||
| QWEN_SANDBOX=docker npx vitest run --root ./integration-tests cli 9>&- | ||
|
|
||
Uh oh!
There was an error while loading. Please reload this page.