Skip to content
Open
Show file tree
Hide file tree
Changes from 5 commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
49 changes: 48 additions & 1 deletion .github/scripts/run-release-docker-integration.sh
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,31 @@ trap 'exit 1' INT TERM
sandbox_revision="$(git rev-parse HEAD)"
sandbox_image="$(node -p "require('./packages/cli/package.json').config.sandboxImageUri")-release-${sandbox_revision}"

# The job-start disk floor gate predates this build: run 37374675168 passed
# it and the runner still died on ENOSPC 24 minutes into this step (#13479).
# Gate the build's own filesystem — the docker data root — at a build-sized
# floor, so a saturated host fails fast with a legible error and a re-run
# lands on an instance with headroom instead of the runner worker crashing
# mid-build. 8 GiB covers a cold builder stage (monorepo install + bundle
# layers) plus the final image with margin. Self-hosted only, like every
# other check-disk-floor.sh call site: an ephemeral hosted runner starts
# with an order of magnitude more free disk than this floor.
check_docker_data_root_floor() {
if [ "$RUNNER_ENVIRONMENT" != 'self-hosted' ]; then
return 0
fi
if [ ! -f .github/scripts/check-disk-floor.sh ]; then
echo "::warning::docker data root floor gate skipped: .github/scripts/check-disk-floor.sh not present at this ref on ${RUNNER_NAME:-this runner}"
return 0
fi
docker_root="$(timeout 60 docker info --format '{{.DockerRootDir}}' 2>/dev/null || true)"
if [ -z "$docker_root" ] || [ ! -d "$docker_root" ]; then
echo "::warning::docker data root floor gate skipped: docker data root '${docker_root:-<unreadable>}' is not a readable directory on ${RUNNER_NAME:-this runner}"
return 0
fi
DISK_FLOOR_MIN_FREE_KB="${DISK_FLOOR_MIN_FREE_KB:-8388608}" bash .github/scripts/check-disk-floor.sh "$docker_root"
Comment thread
qwen-code-dev-bot marked this conversation as resolved.
Outdated
}

if [ "$RUNNER_ENVIRONMENT" = 'self-hosted' ]; then
mkdir -p "${HOME}/.cache/qwen-code-ci"
# Same protocol as e2e.yml: the host daemon lock is held shared for the whole
Expand Down Expand Up @@ -51,7 +76,21 @@ if ! docker image inspect "$sandbox_image" > /dev/null 2>&1; then
exit 1
fi
fi
docker image prune --all --force --filter 'label=org.qwen-code.ci.sandbox=true' --filter 'until=24h' || echo "::warning::old CI sandbox image cleanup failed on ${RUNNER_NAME:-this runner}"
timeout 20m docker image prune --all --force --filter 'label=org.qwen-code.ci.sandbox=true' --filter 'until=24h' || echo "::warning::old CI sandbox image cleanup failed on ${RUNNER_NAME:-this runner}"

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R4-1: The reclaim sequence added in this block is the pool's fourth copy of the same docker reclaim policy, and the four copies now disagree on which commands they run, on whether they are time-bounded, and on --keep-storage. The copy that matters is the peer E2E lane's: .github/scripts/run-e2e-tests.sh:73 runs the identical labelled prune with no timeout at all while holding the very same docker-sandbox-build.lock this script takes at line 77. So the bound added here protects one direction of a two-way contention — a wedged daemon GC on the E2E side still starves this lane's 30-minute wait. Nothing asserts the copies agree either: ci-runner-routing.test.mjs:590-592 pins e2e.yml's step text and ecs-runner/qwen-docker-cleanup.test.mjs:73-77 pins the sweep's --keep-storage 30GB, so a future reclaim change has to be hand-applied to three shell files plus one workflow and drift between them passes CI silently.

The contention premise is not speculative. docs/plans/2026-09-03-test-suite-barrel-cost.md:44 records that the hk4 runners still carry the shared ecs-qwen label, and scripts/tests/lint.test.js:229 records an ecs-qwen-labelled job actually executing on ecs-qwen-hk4-19 two days before this incident.

Witness:

run-e2e-tests.sh:73        docker image prune --all --force --filter 'label=...' --filter 'until=24h'   <- NO timeout, inside flock --wait 1800 7
qwen-docker-cleanup.sh:44  docker image prune --force --filter 'until=24h'                             <- NO timeout
qwen-docker-cleanup.sh:50  timeout 20m docker builder prune --all ... --keep-storage 30GB              <- under flock --nonblock
e2e.yml:421-422            both image prunes                                                            <- NO timeout, under flock --nonblock 9
this diff :83,:87,:96      timeout 20m x3, NO --keep-storage                                            <- the fourth copy

The in-scope minimum is to bound the peer lane's twin — apply the same timeout to run-e2e-tests.sh:73 so both holders of the mutex are bounded. The fuller fix is one shared reclaim helper taking the label filter and an optional reserve, called from all three scripts and the workflow; that is a cross-lane CI-infra refactor and is reasonably a follow-up issue rather than a change here. If neither lands, a comment recording why the release lane is the only one that needs these guards would stop the next reader treating the omission as an oversight.

A shared helper must keep --keep-storage per call site rather than folding the flag sets into one: ecs-runner/qwen-docker-cleanup.sh:50-51 deliberately retains a 30 GB reserve while this line deliberately omits it, because a reserve reclaims nothing on a host holding less cache than the reserve.

If the minimal variant lands, an assertion in scripts/tests/e2e-workflow.test.js — which already reads the E2E script as e2eRunScript at line 13 — that the labelled prune in run-e2e-tests.sh carries a timeout bound is the test to add; nothing pins that today, so it is red before the fix and green after. Please confirm by removing the bound and watching it fail.

中文说明

本块新增的回收序列是这个 runner 池上同一套 docker 回收策略的第四份副本,而四份副本在“执行哪些命令”“是否有时间上限”“是否带 --keep-storage”三个维度上都不一致。最关键的一份是对端 E2E 通道:.github/scripts/run-e2e-tests.sh:73 在持有本脚本第 77 行同一把 docker-sandbox-build.lock 的情况下,运行完全相同的带标签镜像裁剪,却完全没有 timeout。因此这里加的上限只保护了双向争用中的一个方向——E2E 侧守护进程 GC 卡住时,仍然会饿死本通道 30 分钟的等待。也没有任何测试断言这些副本彼此一致:ci-runner-routing.test.mjs:590-592 钉的是 e2e.yml 的步骤文本,ecs-runner/qwen-docker-cleanup.test.mjs:73-77 钉的是日常清理的 --keep-storage 30GB,所以将来任何回收策略变更都要手工同步到三个 shell 文件加一个 workflow,而它们之间的漂移会静默通过 CI。

争用前提并非猜测:docs/plans/2026-09-03-test-suite-barrel-cost.md:44 记录 hk4 的 runner 仍带共享的 ecs-qwen 标签;scripts/tests/lint.test.js:229 记录了一个带 ecs-qwen 标签的作业确实在 ecs-qwen-hk4-19 上执行过,时间就在本次事故前两天。

证据:

run-e2e-tests.sh:73        docker image prune --all --force --filter 'label=...' --filter 'until=24h'   <- 无 timeout,处于 flock --wait 1800 7 之内
qwen-docker-cleanup.sh:44  docker image prune --force --filter 'until=24h'                             <- 无 timeout
qwen-docker-cleanup.sh:50  timeout 20m docker builder prune --all ... --keep-storage 30GB              <- 在 flock --nonblock 之下
e2e.yml:421-422            两个镜像裁剪                                                                 <- 无 timeout,在 flock --nonblock 9 之下
本 diff :83,:87,:96        timeout 20m x3,无 --keep-storage                                            <- 第四份副本

最小改动范围的做法是给对端那份也加上限:对 run-e2e-tests.sh:73 施加同样的 timeout,让互斥锁的两个持有者都受限。更完整的做法是抽出一个共享回收辅助脚本,接受标签过滤器和可选的保留量参数,由三个脚本和 workflow 共同调用;那属于跨通道的 CI 基础设施重构,放到后续 issue 比放在本 PR 更合适。如果两者都不做,建议加一行注释说明为什么只有发布通道需要这些防护,避免后来的读者把这种缺失当成疏漏。

共享辅助脚本必须把 --keep-storage 保留为逐调用点参数,而不是把两组 flag 合并:ecs-runner/qwen-docker-cleanup.sh:50-51 有意保留 30 GB 储备,而本行有意省略它,因为在缓存量小于储备量的主机上,储备参数什么都回收不了。

如果采用最小方案,需要补的测试在 scripts/tests/e2e-workflow.test.js(它在第 13 行已把 E2E 脚本读作 e2eRunScript):断言 run-e2e-tests.sh 中带标签的裁剪携带 timeout 上限。目前没有任何测试钉住这一点,所以修复前它是红的、修复后是绿的。请通过删掉该上限、确认测试变红来验证。

— qwen3.8-max via Qwen Code /review (v0.25.0)

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified before deciding: run-e2e-tests.sh does run the identical labelled prune with no timeout while holding the same docker-sandbox-build.lock this script takes — the one-direction-bound gap is real.

Deferred to the follow-up queue rather than implemented here: bounding the E2E lane means editing the shared CI script every PR's E2E shards run on, which is CI machinery outside this release-lane PR's footprint — the same call your fuller-fix variant makes. The finding, the minimal fix, and the suggested e2e-workflow.test.js pin are recorded in the deferred-findings issue so they survive the merge.

What did land here, as your fallback option: this script now records the asymmetry next to the mutex — every daemon call it makes is time-bounded so a wedged GC cannot starve a peer lane's 30-minute wait, and the E2E lane's unbounded copy of the same prune under the same lock is noted as a cross-lane change tracked separately.

中文说明

决定前已验证:run-e2e-tests.sh 确实在持有本脚本同一把 docker-sandbox-build.lock 的情况下,运行完全相同的带标签镜像裁剪且不带 timeout——“只约束了一个方向”的缺口属实。

本次不直接实现,而是转入后续跟进队列:给 E2E 通道加上限意味着改动每个 PR 的 E2E 分片都在运行的共享 CI 脚本,属于本发布通道 PR 范围之外的 CI 设施——与您给出的完整修复方案(共享回收辅助脚本)同属跨通道改动。该发现、最小修复方案以及建议补在 e2e-workflow.test.js 的钉住测试都已记录在延期发现 issue 中,合并后不会丢失。

按您给出的兜底方案,本脚本已在互斥锁旁记录了这种不对称:本脚本的每次守护进程调用都带时间上限,卡死的 GC 不会饿死对端通道 30 分钟的锁等待;E2E 通道在同一锁下的同款裁剪仍无上限,注释中已注明这是另行跟踪的跨通道改动。

# The labelled prune cannot reach untagged images, and this lane passes
Comment thread
qwen-code-dev-bot marked this conversation as resolved.
Outdated
# --no-prune to the build: an image that went dangling after the daily
# 02:30 UTC sweep would otherwise never be reclaimed.
timeout 20m docker image prune --force --filter 'until=24h' || echo "::warning::dangling image cleanup failed on ${RUNNER_NAME:-this runner}"
Comment thread
qwen-code-dev-bot marked this conversation as resolved.
Outdated
# Image pruning does not reclaim BuildKit's intermediate install/build
# layers. The daily host sweep (ecs-runner/qwen-docker-cleanup) bounds them
# at 30 GB, but this lane builds at the end of the pool's day, hours after
# that sweep. Every daemon call in this branch is bounded so a slow daemon
# GC cannot hold the host build mutex past the E2E lane's 30-minute lock
# wait. No --keep-storage here: it is a reserve, not a quota, so on a host
# with less cache than the reserve — exactly the hosts this line exists
# for — it would reclaim nothing.
timeout 20m docker builder prune --all --force --filter 'until=24h' || echo "::warning::docker build cache cleanup failed on ${RUNNER_NAME:-this runner}"
check_docker_data_root_floor
# See e2e.yml: closing the lock descriptors in the child keeps a descendant
# that outlives this job from holding the lock.
npm run build:sandbox -- -s --no-prune -i "$sandbox_image" 7>&- 8>&- 9>&-
Expand All @@ -67,6 +106,14 @@ if [ "$RUNNER_ENVIRONMENT" = 'self-hosted' ]; then
exec 8>&-
fi

# Run 37374675168 actually died in the vitest phase, ~15 minutes after the

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R4-3: Two point samples now cover three heavy phases, and the phase the incident actually died in is the one nothing measures. This gate is taken once, before two back-to-back vitest docker phases, and cleanup_release_containers is on the EXIT trap, so phase-1 containers and their writable layers are still on the data root when phase 2 starts. The companion instrumentation the repo attaches to every other disk-floor gate is absent from this job: ci.yml runs a 10-second df sampler loop next to its gate and uploads the sample file on failure, and integration_docker in release.yml has neither.

A host that clears this sample with just over 8 GiB free can cross zero during the cli phase — container writable layers, transcripts and the runner's own diagnostic log all land on that filesystem — and the worker dies on ENOSPC mid-phase with no step conclusion, so even the always() container-cleanup step never runs and the auto-filed issue again carries no actionable signal. That is the identical shape as run 37374675168. This is not covered by the PR's out-of-scope statement: in-lane instrumentation is neither an end-to-end lane validation nor a host-level capacity control, and #10394 already established it as the companion to every disk-floor gate in ci.yml.

Witness:

DFSAMPLE occurrences:  .github/workflows/ci.yml = 6 (L636, L765, L780, L1158, L1708, L1890)
                       .github/workflows/release.yml = 0
sed -n '530,585p' release.yml | grep -E 'DFSAMPLE|df -|upload-artifact|if: .*failure'  ->  NONE FOUND
only upload-artifact in release.yml is at :340, owned by quality_build
  (job headers: integration_docker = 530, audio_capture_prebuilds = 583)
ci.yml:632-643 is the pattern: sample_disk + ( while sleep 10; do sample_disk; done ) & + trap
  dual sink at :639-640
script:21  trap cleanup_release_containers EXIT     gate at :119, vitest at :122 and :123

Attach the sampler #10394 established to this lane: start a 10-second df loop over "$docker_root" and "${RUNNER_TEMP}" before the first npx vitest, stop it from the existing trap, and add an if: failure() actions/upload-artifact step for the sample file to integration_docker. A third gate sample between the two vitest invocations would also help, but the sampler is what turns a recurrence into a diagnosable one.

Each sample must go to the job log as well as to the file — ci.yml:780 does exactly echo "$sample"; echo "$sample" >> "$DISK_SAMPLES" 2>/dev/null || true — because in the incident the worker itself crashed, the step got no conclusion and even the always() step never ran, so a file plus an if: failure() upload alone would have been lost the same way.

A new case in scripts/tests/release-workflow.test.js in the style of the existing pin at line 349 (gates every pool-routed job on a disk floor before its heavy steps) is the test to add: assert releaseYaml.jobs.integration_docker.steps contains a step whose run starts the sampler loop and an if: failure() upload-artifact step whose path is the sample file. Deleting either must redden it — please confirm by removing one and re-running.

中文说明

现在用两个时间点采样去覆盖三个重负载阶段,而事故真正丧生的那个阶段恰恰没有任何测量。这个门只取一次样,位于两个连续执行的 vitest docker 阶段之前;而 cleanup_release_containers 挂在 EXIT trap 上,所以第一阶段产生的容器及其可写层在第二阶段开始时仍留在数据根目录上。仓库为其它每一个磁盘下限门配备的伴生观测手段在本作业中完全缺席:ci.yml 在其门旁边运行每 10 秒一次的 df 采样循环,并在失败时上传采样文件,而 release.yml 的 integration_docker 两者都没有。

一台以略高于 8 GiB 的余量通过本次采样的主机,可能在 cli 阶段期间归零——容器可写层、测试记录以及 runner 自身的诊断日志都写在同一个文件系统上——于是 worker 在阶段中途因 ENOSPC 死亡,步骤没有任何结论,连 always() 的容器清理步骤都不会运行,自动提交的 issue 再次不带任何可操作信号。这与运行 37374675168 的形态完全相同。这不在本 PR 的“超出范围”声明覆盖之内:通道内观测既不是端到端通道验证,也不是主机级容量控制,而且 #10394 已经把它确立为 ci.yml 中每个磁盘下限门的伴生措施。

证据:

DFSAMPLE 出现次数:  .github/workflows/ci.yml = 6(L636、L765、L780、L1158、L1708、L1890)
                     .github/workflows/release.yml = 0
sed -n '530,585p' release.yml | grep -E 'DFSAMPLE|df -|upload-artifact|if: .*failure'  ->  NONE FOUND
release.yml 中唯一的 upload-artifact 在 :340,属于 quality_build
  (作业起始行:integration_docker = 530,audio_capture_prebuilds = 583)
ci.yml:632-643 即该模式:sample_disk + ( while sleep 10; do sample_disk; done ) & + trap
  双写形式在 :639-640
script:21  trap cleanup_release_containers EXIT     门在 :119,vitest 在 :122 与 :123

请把 #10394 确立的采样器接到本通道上:在第一个 npx vitest 之前启动一个每 10 秒对 "$docker_root" 和 "${RUNNER_TEMP}" 取样的 df 循环,由现有 trap 停止它,并为 integration_docker 增加一个 if: failure() 的 actions/upload-artifact 步骤上传采样文件。在两个 vitest 调用之间再加一次门采样也有帮助,但真正把“再次发生”变成“可诊断”的是采样器。

每个采样必须同时写入作业日志和文件——ci.yml:780 正是 echo "$sample"; echo "$sample" >> "$DISK_SAMPLES" 2>/dev/null || true——因为事故中 worker 自身崩溃、步骤没有结论、连 always() 步骤都未运行,所以只有文件加 if: failure() 上传同样会丢失。

需要补的测试在 scripts/tests/release-workflow.test.js,可参照第 349 行已有的钉法(gates every pool-routed job on a disk floor before its heavy steps):断言 releaseYaml.jobs.integration_docker.steps 中包含一个 run 启动采样循环的步骤,以及一个 if: failure() 的 upload-artifact 步骤且其 path 为采样文件。删除其中任一项都必须让测试变红——请通过删掉一项并重跑来确认。

— qwen3.8-max via Qwen Code /review (v0.25.0)

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified before deciding: release.yml's integration_docker job has no df sampler and no failure upload — ci.yml carries six DFSAMPLE loops, this workflow has none, and its only upload-artifact step belongs to quality_build. The gap is real and matches the incident shape (worker died mid-phase, no step conclusion, even the always() cleanup never ran).

Deferred to the follow-up queue rather than implemented here: attaching the #10394 sampler means adding steps to a workflow this PR never touches — CI machinery outside this PR's footprint. The full design you specified (10-second loop over the docker root and RUNNER_TEMP, dual job-log/file sink per ci.yml:780, if: failure() upload-artifact step, plus the new pin in the style of the line-349 test) is recorded in the deferred-findings issue so it survives the merge.

The secondary ask — a third gate sample between the two vitest invocations — was considered and not added: a point sample there does not diagnose the incident shape (only the streamed sampler survives a worker crash), and each added gate is flake surface on the lane. It rides along in the same deferral record.

中文说明

决定前已验证:release.yml 的 integration_docker 作业确实没有 df 采样器、也没有失败上传——ci.yml 有六处 DFSAMPLE 循环,本工作流一处没有,唯一的 upload-artifact 步骤属于 quality_build。缺口属实,且与事故形态吻合(worker 在阶段中途死亡,步骤没有结论,连 always() 清理步骤都未运行)。

本次不直接实现,而是转入后续跟进队列:接入 #10394 确立的采样器需要向本 PR 从未改动过的工作流添加步骤,属于本 PR 范围之外的 CI 设施。您指定的完整设计(对 docker 根目录与 RUNNER_TEMP 的 10 秒循环、按 ci.yml:780 同时写作业日志与文件的双写、if: failure() 的 upload-artifact 步骤,以及仿照第 349 行钉法的新测试)已记录在延期发现 issue 中,合并后不会丢失。

次要建议——在两个 vitest 调用之间再加一次门采样——经考虑后未采纳:在那里做一次点采样无法诊断事故形态(worker 崩溃时只有流式采样能留存),而且每增加一道门都会给通道带来抖动面。该点已一并记入同一条延期记录。

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R4-3: Still stands. Two point samples now cover three heavy phases, and the phase the incident actually died in is the one nothing measures. This gate is taken once, before two back-to-back vitest docker phases, and cleanup_release_containers is on the EXIT trap, so phase-1 containers and their writable layers are still on the data root when phase 2 starts. The companion instrumentation the repo attaches to every other disk-floor gate is still absent from this job: ci.yml runs a 10-second df sampler loop next to its gate and uploads the sample file on failure, and integration_docker in release.yml has neither.

A host that clears this sample with just over 2 GiB free can cross zero during the cli phase — container writable layers, transcripts and the runner's own diagnostic log all land on that filesystem — and the worker dies on ENOSPC mid-phase with no step conclusion, so even the always() container-cleanup step never runs and the auto-filed issue again carries no actionable signal. That is the identical shape as run 37374675168, which is what this PR exists to make diagnosable. A point sample cannot survive a worker crash; only a streamed sampler can.

Witness (measured in the review worktree at head 46c54cfbe2):

grep -c DFSAMPLE .github/workflows/ci.yml        -> 6
grep -c DFSAMPLE .github/workflows/release.yml   -> 0
grep -c 'while sleep 10' .github/workflows/release.yml -> 0
upload-artifact steps inside integration_docker  -> 0
   (release.yml's single upload-artifact belongs to quality_build)

You verified this gap in comment 4211178534 and deferred it to the follow-up queue because the remedy adds steps to a workflow this PR never touches, which is a reasonable scoping call and is recorded in the deferred-findings issue. Nothing landed this round, so the entry carries forward rather than closing — the deferral record is now the only thing carrying it. The design you specified there is the fix: a 10-second df loop over "$docker_root" and "${RUNNER_TEMP}" started before the first npx vitest, stopped from the existing trap, plus an if: failure() actions/upload-artifact step for the sample file on integration_docker.

Each sample must go to the job log as well as to the file, which is what ci.yml:780 does — echo "$sample"; echo "$sample" >> "$DISK_SAMPLES" 2>/dev/null || true — because in the incident the worker itself crashed, the step got no conclusion and even the always() step never ran, so a file plus an if: failure() upload alone would have been lost the same way. A new case in the style of the existing pin at release-workflow.test.js:349 (gates every pool-routed job on a disk floor before its heavy steps) is the test to add: assert releaseYaml.jobs.integration_docker.steps contains a step whose run starts the sampler loop and an if: failure() upload-artifact step whose path is the sample file, then delete each in turn and confirm it reds.

中文说明

[建议] R4-3:仍然存在。现在是两次点采样覆盖三个重载阶段,而事故真正死掉的那个阶段没有任何东西在测量。这道门只取一次样,位于两个连续的 vitest docker 阶段之前,而 cleanup_release_containers 挂在 EXIT trap 上,因此阶段一的容器及其可写层在阶段二开始时仍然在数据根目录上。仓库为其他每一道磁盘下限门配备的伴生指标采集,在本作业仍然缺席:ci.yml 在其门旁跑一个 10 秒的 df 采样循环并在失败时上传采样文件,而 release.yml 的 integration_docker 两者都没有。

一台以刚好超过 2 GiB 的剩余空间通过本次采样的主机,可能在 cli 阶段降到零——容器可写层、日志输出以及 runner 自身的诊断日志都写在那个文件系统上——于是 worker 在阶段中途因 ENOSPC 死掉,步骤没有任何结论,连 always() 的容器清理步骤都不会运行,自动提交的 issue 再次不带任何可操作信号。这与运行 37374675168 的形态完全一致,而让它可诊断正是本 PR 的目的。点采样无法在 worker 崩溃后存活,只有流式采样可以。

证据(在 head 46c54cfbe2 的评审 worktree 中实测):

grep -c DFSAMPLE .github/workflows/ci.yml        -> 6
grep -c DFSAMPLE .github/workflows/release.yml   -> 0
grep -c 'while sleep 10' .github/workflows/release.yml -> 0
integration_docker 内的 upload-artifact 步骤  -> 0
   (release.yml 唯一的 upload-artifact 属于 quality_build)

你已在评论 4211178534 中验证了该缺口并转入后续跟进队列,理由是补救措施需要向本 PR 从未改动的工作流添加步骤——这是合理的范围判断,也已记入延期发现 issue。本轮没有任何改动落地,因此该条目向前携带而不是关闭——现在延期记录是唯一携带它的东西。你在那里指定的设计就是修复方案:在第一次 npx vitest 之前启动一个针对 "$docker_root" 与 "${RUNNER_TEMP}" 的 10 秒 df 循环,由现有 trap 停止,并为 integration_docker 添加一个 if: failure() 的 actions/upload-artifact 步骤上传采样文件。

每次采样必须同时写入作业日志与文件,也就是 ci.yml:780 的做法——echo "$sample"; echo "$sample" >> "$DISK_SAMPLES" 2>/dev/null || true——因为在那次事故中 worker 本身崩溃、步骤没有结论、连 always() 步骤都未运行,所以仅靠文件加 if: failure() 上传会以同样的方式丢失。要补的测试可仿照 release-workflow.test.js:349 现有的钉法(gates every pool-routed job on a disk floor before its heavy steps):断言 releaseYaml.jobs.integration_docker.steps 包含一个 run 启动采样循环的步骤,以及一个 if: failure() 且 path 为采样文件的 upload-artifact 步骤,然后逐个删除并确认它变红。

— qwen3.8-max via Qwen Code /review (v0.25.0)

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified and still deferred to the follow-up queue: the remedy (a 10-second df sampler over the docker data root and RUNNER_TEMP started before the first npx vitest, stopped from the existing EXIT trap, each sample echoed to the job log as well as the file — the ci.yml:780 pattern, because the incident's worker crash loses a file-only artifact — plus an if: failure() upload-artifact step) adds steps to .github/workflows/release.yml, which this PR never touches. Expanding into a new workflow file mid-review is exactly the scope drift the footprint boundary exists to prevent. The deferral is recorded in the deferred-findings issue so it survives the merge; the design above is the fix to schedule from there.

中文说明

已核实,仍延后至跟进队列:补救方案(在第一次 npx vitest 之前启动一个针对 docker 数据根目录与 RUNNER_TEMP 的 10 秒 df 采样循环,由现有 EXIT trap 停止,每条采样同时写入作业日志与文件——即 ci.yml:780 的做法,因为事故中 worker 崩溃会丢失仅存文件的产物——外加一个 if: failure() 的 upload-artifact 步骤)需要向 .github/workflows/release.yml 添加步骤,而本 PR 从未改动该文件。在评审中途扩展到新的工作流文件,正是足迹边界要防止的范围漂移。该延后已记录在延期发现 issue 中,合并后依然保留;上述方案就是应从那里排期的修复。

# build finished: re-gate now that the build's peak and the image it leaves
# behind have landed on this filesystem. The gate sits outside the
# image-missing branch so a re-run that finds the image already cached on a
Comment thread
qwen-code-dev-bot marked this conversation as resolved.
# still-saturated host is gated before vitest writes container layers to the
# same filesystem.
check_docker_data_root_floor
Comment thread
qwen-code-dev-bot marked this conversation as resolved.
Outdated

# The package.json docker test scripts each rebuild the sandbox image. Run
# vitest directly here so this job reuses the image built above.
QWEN_SANDBOX=docker npx vitest run --root ./integration-tests cli 9>&-
Expand Down
Loading
Loading