Skip to content

test(integration): make the utf-bom defaultFileEncoding test deterministic - #13196

Merged
wenshao merged 1 commit into
QwenLM:mainfrom
Bug-killer-prog:utf-bom-encoding-deterministic
Oct 5, 2026
Merged

wenshao merged 1 commit into
QwenLM:mainfrom
Bug-killer-prog:utf-bom-encoding-deterministic

Conversation

@Bug-killer-prog

Copy link
Copy Markdown
Contributor

What this PR does

Converts the last real-model test in the BOM integration suite (should create new file with BOM when defaultFileEncoding is utf-8-bom) to the deterministic forced-tool-call pattern used by the seven sibling tests in the same file: the fake server now scripts the write_file call, while the real tool still writes the real file, so the byte-level assertions are unchanged.

Why it's needed

The defaultFileEncoding case was the only test in utf-bom-encoding.test.ts that drove an actual model round-trip (rig.run + waitForToolCall, up to 300s per attempt). On the loaded Docker self-hosted runner that round-trip exceeded the vitest timeout and failed the integration_docker job, blocking the v0.24.8-preview.0 release (#13066). The issue discussion diagnoses this as a test determinism problem and suggests converting this test to runForcedToolCallScenario with a fake write_file call — this PR applies exactly that.

Reviewer Test Plan

How to verify

npm run build && npm run bundle, then cd integration-tests && cross-env QWEN_SANDBOX=false npx vitest run cli/utf-bom-encoding.test.ts.

Expected: the defaultFileEncoding case now finishes in seconds instead of waiting on a model round-trip, and still asserts the EF BB BF prefix written by the real write_file tool under the defaultFileEncoding: utf-8-bom setting. No real API key is involved (fake OpenAI server).

Evidence (Before & After)

N/A — no user-facing change (test-only).

Tested on

OS Status
🍏 macOS ⚠️
🪟 Windows ✅
🐧 Linux ⚠️

Verified locally on Windows 11 (Node v24.12.0): the converted test passes in ~6s; the other seven cases stay skipped on Windows by pre-existing design. Linux/macOS lanes are covered by CI.

Environment

Local bundle (npm run bundle), QWEN_SANDBOX=false, fake OpenAI server — no real model credentials.

Risk & Scope

  • Main risk or tradeoff: the scripted tool call no longer exercises "the model decides to create a file" — the same tradeoff the seven sibling tests already accept; the chain under test (settings → real write_file tool → BOM bytes) is untouched, and unit coverage of the BOM logic remains in place.
  • Not validated / out of scope: the two BOM-preservation cases in the first describe block use the same real-model pattern but were not the release blocker, and that block is Windows-skipped so they cannot be verified from a Windows dev machine — left for a follow-up.
  • Breaking changes / migration notes: none (test-only change).

Linked Issues

Closes #13066

中文说明

这个 PR 做了什么

把 BOM 集成测试套件中最后一个走真实模型的用例(should create new file with BOM when defaultFileEncoding is utf-8-bom)改为与同文件其余七个用例一致的确定性 forced-tool-call 模式:由假服务器脚本化下发 write_file 调用,真实的 write_file 工具仍然写真实文件,因此字节级断言不变。

为什么需要

defaultFileEncoding 这个用例原本是 utf-bom-encoding.test.ts 中唯一走真实模型往返的测试(rig.run + waitForToolCall,每次最长等 300s)。在高负载的 Docker 自托管 runner 上,该往返超过 vitest 超时,导致 integration_docker job 失败并阻塞 v0.24.8-preview.0 发布(#13066)。issue 讨论中已诊断这是测试确定性问题,并建议把该用例改为 runForcedToolCallScenario + 假 write_file 调用——本 PR 即按此方案实现。

评审验证计划

如何验证

先 npm run build && npm run bundle,再 cd integration-tests && cross-env QWEN_SANDBOX=false npx vitest run cli/utf-bom-encoding.test.ts。预期:转换后的用例几秒内完成,不再等待模型往返,仍断言真实 write_file 工具在 defaultFileEncoding: utf-8-bom 设置下写出的 EF BB BF 前缀。无需真实 API key(假 OpenAI 服务器)。

前后对比证据

N/A —— 无用户可见变更(仅测试)。

测试环境

OS 状态
🍏 macOS ⚠️
🪟 Windows ✅
🐧 Linux ⚠️

已在 Windows 11(Node v24.12.0)本地验证:转换后的用例约 6 秒通过;其余七个用例在 Windows 上按既有设计跳过;Linux/macOS 由 CI 覆盖。

风险与范围

  • 主要风险/取舍:脚本化的工具调用不再覆盖"模型决定创建文件"这一环——同文件其余用例本就如此取舍;被测链路(设置 → 真实 write_file 工具 → BOM 字节)未变,BOM 逻辑的单元覆盖仍在。
  • 未验证/超出范围:第一个 describe 块中的两个 BOM 保留用例使用相同的真实模型模式,但它们不是本次发布阻塞点;该块在 Windows 上被跳过,无法在 Windows 开发机上验证,留作后续跟进。
  • 破坏性变更:无(仅测试)。

关联 Issue

Closes #13066

…istic

The 'should create new file with BOM when defaultFileEncoding is utf-8-bom'
integration test was the only case in utf-bom-encoding.test.ts that drove a
real model round-trip (rig.run + waitForToolCall, up to 300s per attempt).
On the loaded Docker self-hosted runner that round-trip can exceed the
vitest timeout, which blocked the v0.24.8-preview.0 release (QwenLM#13066).

Convert it to runForcedToolCallScenario with a fake write_file call, the
same deterministic pattern as the sibling tests in this file. The real
write_file tool still executes, so the BOM assertions are unchanged. Also
unstub the env vars the scenario stubs, matching the first describe block.

Closes QwenLM#13066

Co-authored-by: doudouOUC <[email protected]>
@Bug-killer-prog

Copy link
Copy Markdown
Contributor Author

Head 319c941 is fully green now — every check has completed with 0 failures, nothing running. The earlier "Downgraded from Approve to Comment: CI still running" looks like it raced a check that was still pending at that moment. Could a maintainer re-run @qwen-code /triage so the approval can be re-posted? (Review ledger for the downgrade shows "findings":[] — no outstanding code issues.)

@wenshao

wenshao commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

Verdict: merge-ready — local maintainer verification, 26 verification cells executed, 0 unexpected failures. Verified head 319c941f3d05b36990ef9d411515948fbf76fde8 against base 47463b79a7dcf1559d03fd06611e39c7a2dec7d5 (= head parent; single-commit PR).

中文摘要

结论:merge-ready(从验证角度可以合并)。 在 macOS(darwin/arm64,Node 22.23.1)本地真实构建环境中完成验证,head/base 双 worktree 各自独立 pnpm install --frozen-lockfile + npm run build && npm run bundle。

  • A/B 证明(核心声明成立):同一条被转换的测试,在完全相同的无凭据环境(空 HOME、无任何 OPENAI_*/QWEN_* 变量)下——base 红(No auth type is selected,三次尝试同一签名),head 绿且约 2 秒完成。head 侧不再需要任何真实模型。base 的发布阻塞签名(release run 36602108252:Test timed out in 300000ms × 3,单用例 900s)与本地表现一致:该测试本质上依赖一次真实模型往返。
  • 全文件对照:整个 utf-bom-encoding.test.ts(8 用例)两臂同环境运行:base 5 绿 3 红 → head 6 绿 2 红,差值恰好是被转换的这一个用例;剩余 2 红是两个既有真实模型用例,本 PR 未触碰,两臂表现一致。
  • 变异矩阵(测试非空转):M1 把设置改为 utf-8(无 BOM)→ 红(expected 99 to be 239,99='c' 即真实 write_file 未写 BOM);M2 把断言期望值改错 → 红(expected 239 to be +0,文件确实以 0xEF 开头)。两个变异体全部被杀,证明测试仍完整覆盖「设置 → 真实 write_file 工具 → 磁盘 EF BB BF 字节」链路。每次变异后恢复并确认 git status 干净。
  • 门禁:tsc -p integration-tests 通过;eslint 零告警(先植入未使用变量确认 eslint 活性,再移除复验);prettier 通过。
  • 合并状态:base 之后 main 未触碰该文件,PR MERGEABLE(BLOCKED 仅为评审门禁)。
  • 未覆盖:Docker 沙箱车道(本机 macOS 无法复现该 lane,CI 覆盖);两个既有真实模型用例(PR 作者已声明留作后续);Windows 实跑(该文件在 Windows 按既有设计跳过)。

Central claim

The defaultFileEncoding case — the only real-model round-trip left in utf-bom-encoding.test.ts, and the test whose 3 × 300s timeouts blocked the v0.24.8-preview.0 release (#13066) — now runs deterministically against a fake OpenAI server in seconds, while the real write_file tool still writes the file, so the byte-level BOM assertions are unchanged.

A/B load-bearing proof

Two worktrees that differ only in the PR commit, each independently installed (--frozen-lockfile) and built (npm run build && npm run bundle). Same targeted command on both, in an identical credential-free environment (empty HOME, no OPENAI_*/QWEN_* variables):

Cell Tree Result
Head 319c941f3d PASS in ~2s — no model involved
Control 47463b79a7 FAIL (3 attempts, retry x2) — No auth type is selected. Please configure an auth type ... before running in non-interactive mode

The base arm's red is the bug's own signature in local form: the old test cannot pass at all without a live authenticated model, and with one it costs a real round trip per attempt — which is exactly what exceeded the 300s vitest timeout three times in a row on the loaded Docker release runner (quoted from release run 36602108252):

× BOM with defaultFileEncoding configuration > should create new file with BOM
  when defaultFileEncoding is utf-8-bom  900131ms (retry x2)
→ Test timed out in 300000ms.

Head: converted test green in seconds, credential-free

Base: same test red without a live model

Control cleanliness: readlink of node_modules/@qwen-code/qwen-code-core resolves into each tree's own packages/core on both sides — no cross-tree leak. (Head was additionally run with a scrubbed HOME — no ~/.qwen/settings.json — and still passes, since runForcedToolCallScenario passes explicit CLI flags that outrank user settings.)

Full-file parity (collateral sweep)

The whole file (8 tests) run on both arms under the same credential-free environment:

Arm Result
Base 47463b79a7 5 passed / 3 failed
Head 319c941f3d 6 passed / 2 failed

The red→green delta is exactly the converted test. The 2 remaining reds are the pre-existing real-model preservation cases (should preserve UTF-8 BOM when editing... / ...when overwriting...), untouched by this PR, failing identically on both arms with the same no-auth signature (the author explicitly defers them to a follow-up; they are Windows-skipped by design and were not the release blocker).

Full-file A/B parity

Mutation matrix (vacuity check)

Two mutants applied to the converted test at the head tree, one at a time; expected outcome for both: killed (red).

Mutant Rationale Result
M1: setting utf-8-bom → utf-8 real write_file then emits no BOM KILLED — expected 99 to be 239 (99 = c, the file really starts with content, not BOM)
M2: expectation 0xef → 0x00 corrupted assertion KILLED — expected 239 to be +0 (the file really does start with 0xEF)

M1 is the load-bearing one: it proves the converted test still exercises the full chain settings.defaultFileEncoding → real write_file tool → EF BB BF bytes on disk, i.e. the scripted fake write_file call does not make the test vacuous. After each mutant the file was restored, git status verified clean, and the pristine rerun is green (that rerun is image 01).

Mutation matrix: both mutants killed

Targeted gates (head 319c941f3d)

Gate Result
Converted test, -t targeted, credential-free PASS ~2s (run twice: plain + capture)
Full file, credential-free 6/8 pass; 2 pre-existing real-model reds (parity with base)
tsc -p integration-tests/tsconfig.json exit 0
eslint integration-tests/cli/utf-bom-encoding.test.ts 0 problems
prettier --check on the file pass

ESLint liveness was proven before citing it: a planted const UNUSED_PROBE = 42; was reported (@typescript-eslint/no-unused-vars), then removed and the file re-verified clean.

Findings

None. One observation, not a code issue: the PR also adds afterEach(() => vi.unstubAllEnvs()) to the second describe block, matching the first block — this matters because runForcedToolCallScenario stubs env vars, and the new test is that block's only vi.stubEnv consumer; the addition is correct and consistent.

Not covered

  • Docker sandbox lane — this round ran on macOS (darwin/arm64) with QWEN_SANDBOX=false; the release-blocking lane is Linux/Docker. The mechanism under test (model round trip vs scripted fake call) is platform-independent, and the PR's own CI (incl. Integration Tests (no-AK, No Sandbox)) is green.
  • The two pre-existing real-model preservation cases — untouched by this PR; verified only for cross-arm parity (identical no-auth failure on both arms).
  • Windows — the whole file is Windows-skipped by pre-existing design.

Methodology

macOS (darwin/arm64), Node v22.23.1. Head tree: git worktree at 319c941f3d; base tree at 47463b79a7 (head's parent). Each tree: corepack pnpm install --frozen-lockfile → npm run build → npm run bundle. Test env: env -i HOME=<empty tmp> PATH=$PATH QWEN_SANDBOX=false CI=true (credential-free; CI=true only extends waitForToolCall polling, never reached on red arms since the CLI exits immediately). Commands: npx vitest run --root ./integration-tests cli/utf-bom-encoding.test.ts [-t ...]. Mutants: edit at head tree, rerun, git checkout -- restore, git status verify. Counts: 2 targeted head runs + 2 targeted base runs (expected-red) + 8 full-file head cells + 8 full-file base cells + 2 mutants + 4 gates (eslint probe, eslint clean, prettier, tsc) + 2 workspace-link realpath checks = 26 pass, 0 unexpected fail.

— local verification by @wenshao

@wenshao
wenshao added this pull request to the merge queue Oct 5, 2026
Merged via the queue into QwenLM:main with commit 41f9711 Oct 5, 2026
78 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Release Failed for v0.24.8-preview.0 on 2026-09-29

2 participants