Repository navigation
feat(managed-agent): let creators change a bound Session's directory (W2) - #13247
Conversation
…(W2) Implement the W2 slice of #12380: a durable, idempotent cwd change for a Workspace-bound Session under the workspace-files opt-in. * Admission replays by idempotency key before any revision or busy check, is creator-only (unreadable 404 / readable non-creator 403 session_operation_forbidden, matching the merged bound lifecycle), and refuses with typed 400/409 codes in a pinned order. * Settlement probes the target against the administrator mounts (canonical path, file identity, storage guard, the exact directory rule a tool turn enforces) and commits one transaction: binding update with a revision CAS, operation completion or a terminal failure_code that never retries, and one session.context.changed event per completed change. * The next tool turn installs the committed binding on a fresh Runtime Session, so no worker or Harness protocol change is needed. * Contract v1.29.0 -> v1.30.0 flips both routes and their schemas to implemented; Flyway V34 adds the operation's cwd columns.
E2E test report — W2 controlled same-Workspace cwd changePlan + full notes: Baseline (pre-change, earlier in the slice)
Results (this PR's head)
Review notes and deferrals
MySQL lane: not run locally — CI runs the hosted MySQL suites; the H2 coverage uses |
…rade paths The V34 cwd columns are additive, but the operation mapper read them unconditionally, so any read against a schema pinned before V34 - exactly what the MariaDB retention upgrade IT constructs - failed with bad SQL grammar. Map them through a metadata existence check instead (absent becomes null, which no operation row can carry anyway), and pin the invariant with an H2 twin of the upgrade test.
|
CI note — the first run's Root cause: Fix: additive reads go through a metadata existence check — an absent column maps to null, content no operation row can carry anyway — with the invariant pinned locally by a new H2 twin ( |
Real-stack verification: W2 same-Workspace cwd change @
|
| Scenario | Result |
|---|---|
S1: both surfaces 202 → completed within 0.12 s; normalized replay (child2//, child2/./, ./child2) gives the same id with replayed=true; same key with a different target or revision → 409 idempotency_conflict; stale revision → 409 context_revision_conflict; cross-surface operation reads; DB row COMPLETED/CONFIRMED + receipt; exactly one session.context.changed per change on public + WebShell live SSE |
24/24 |
S2: 401 actor_required; 400 key/body/type checks; 12 lexical shapes → invalid_cwd before any Session fact; unknown / cross-tenant / stranger → 404; reader and other Workspace creator → 403 session_operation_forbidden (also with a stale revision, so no state leak); revoked can_create / DRAINING / generation drift → 409 workspace_unavailable; active initial Turn → 409 session_context_busy (CAS and 403 take precedence); no operation row on any refusal |
50 ✓ |
S2b: CLOSING/CLOSED/ARCHIVING/ARCHIVED/DELETING → 409 session_state_conflict, DELETED → 404; replay of a completed change still answers before the gate. Status was seeded by SQL, because bound close is unavailable on the macOS stack. |
14/14 |
S3: missing, file, FIFO, symlink inside / outside, inner symlink, case alias, NFD alias, component longer than NAME_MAX → failed/workspace_unavailable with binding, revision and events unchanged and no retry; NFC, CJK + spaces, nested, ., same directory → completed, revision +1 |
16/16 |
| S4: 10 rounds × 20 concurrent admissions (public and WebShell interleaved) → exactly one 202 / commit / event per round; 20 concurrent same-key requests with mixed spellings → one row; same key with different payloads → 409 | 3/3 |
S5: injected commit failure ×2 → retried, attempt_count=2, one commit; kill -9 inside the commit (after the probe) → reads stay on the old pair and the op reads installing, then the op is reclaimed 59.6 s later (lease) at claim generation 2 with one event; kill -9 inside the claim → completes 8 s after the kill; revoked can_create / can_read, DRAINING, generation bump, directory removed or swapped between claim and commit → failed, binding kept, a later change still works |
11/11¹ |
| S7: base → routes 404; upgrade V33 → V34; a pre-upgrade Session changes directory; rollback/roll-forward as in N2 | 8/8 |
S9/S13 on head ⊕ main 2b15eac862: the next later Turn writes into child2/ (not child/), into the old directory after a refused change, and at the root after .; a directory replaced by a symlink to outside after the commit → next Turn fails hosted_turn_failed and nothing is written outside; a Turn racing an open change behaves as in N5 |
6/6 + 1/1 |
Unit (H2) / HostedPublicWorkspaceIT (H2, MySQL 8.4.7) / checkstyle |
485/485 · 3/3 · 3/3 (+1 Linux-only skipped) · clean |
| Mutation: 23 mutants of the admission order, CAS, busy barrier, creator/Registry checks, commit re-checks, lease guards, probe rules, event and kind branch, run against the PR's focused tests | 19/23 killed; survivors M15/M18/M18b explained above; M17 (drop only the lease-expiry predicate of failCwdChangeOperation; owner/generation still guard) is near-equivalent |
¹ D1 (can_read revoked) was confirmed from the operation row (FAILED / workspace_unavailable). The probe polled as the revoked actor and, correctly, got 404.
Evidence (rig, probes, MySQL fault/hold triggers, per-arm results, mutation ledger, candidate patch): wenshao/qwen-code@84e0818 → pr13247/
中文版
真实栈验证:W2 同 Workspace 目录变更 @ a2d7bd3b02
结论:机制在真实栈上成立。合并前需要:基于最新 main 重新合并、迁移改号(B1)、修复结算探针(F1)。 F1 原本处于潜伏状态,直到今天 01:42 UTC #13112 合入 main,「目录不可读」这种情形现在在 main 上可以直接触达。
环境:
- macOS 宿主、JDK 21、MySQL 8.4.7。
- 真实的 Spring fat jar(内嵌 Runtime Broker)、打包的 Hosted Harness(用本 head 构建的
dist/cli.js)、确定性的 fixture 模型、rig 用的 actor 适配器。全部走 HTTP,事实从 MySQL 读回核对。 - 对照臂:
head- 合并基
a011f669 head ⊕ main fa795e0232head ⊕ main 2b15eac862:feat(managed-agent): let a Workspace-bound Session's creator submit, cancel and rename #13112 合入时的提交是d20a1895,与我之前做的试合并是同一个提交。- 以上每一臂再各加一份候选补丁。
合并事项
B1 — Flyway V34 与 main 撞号(阻塞,机械性)。 #13225(00:17 UTC 合入)新增了 V34__managed_tool_output_collection.sql。
- 现象:与
fa795e0232合并时 git 没有冲突,但服务启动即失败,报Found more than one migration with version 34,一条迁移都没有执行。 - CI 为何是绿的:它跑在旧的 merge ref 上。
- feat(managed-agent): let a Workspace-bound Session's creator submit, cancel and rename #13112 合入后,
README.md和 OpenAPI 的info.version(1.30.0 vs main 的 1.29.0)也出现了文本冲突,GitHub 现在显示CONFLICTING。 - 已验证的修法:把文件改名为
V35__managed_cwd_operation.sql即可。- main jar 建出的库(V34)能升级到 V35。
- main jar 创建的会话在升级后可以变更目录。
- S1 24/24 通过;
head ⊕ 2b15eac862同样通过,该树上单测 553/553、后续 Turn 6/6。
- 需要协调:开放中的 feat(managed-agent): broker authentication and broker-provisioned writer credentials #13210(V34)和 fix(managed-agent): stop database amplification on session hot paths #13217(V34、V35)也占用了这些编号。
F1 — 结算探针会接受 worker 一定会拒绝的目录(应修)。
- 原因:
WorkspaceRuntimeResolver.requireDirectory检查了包含关系、isDirectory(NOFOLLOW_LINKS)和 realpath 一致,但没有访问权限检查;worker 有这一步(managed-context-worker.ts:123调用fs.access(directory, R_OK | X_OK))。 - 设计文档点出了正是这个差异,并称它会「在首个 Turn 暴露、仍带类型、绝不错投目录」,bot 评审也是同样推理。「不错投」成立,「带类型」不成立。
- 在
head ⊕ main 2b15eac862上,每次试验用一块全新的存储,目录权限分别为 000、444、111(各 3/3):- 目录变更报
completed。 - 下一个后续 Turn 认领存储后安装失败,Harness 日志为
recovery blocked: … HTTP 503 (runtime_session_acquire_failed)。qwen_runtime_session那一行停在ACQUIRING,存储执行租约一直被占着,所以 Turn 一直处于RUNNING。 - 同一存储上的其他会话,工具 Turn 全部失败:
409 workspace_busy→hosted_turn_failed,包括同一 Workspace 里其他创建者的会话。 - 该会话自身再变更目录被拒(
409 session_context_busy),后续 Turn 得到turn_active。 - 之后
chmod 755、重启 Spring 与 Harness 都没能解除:又过了 3 分钟,Turn 仍是 RUNNING,租约仍被占用。这是在 macOS 非 durable 栈上的结果,Linux durable 的恢复路径我没有跑。 - 以 root 运行的 worker 不受权限位约束,但遇到其他任何访问拒绝时会走到同一条路径。
- 目录变更报
- 归因:创建(G0)时直接指向 000 目录,同样会卡死(不涉及 W2),所以 Turn 路径上的卡死在 main 上本来就存在。W2 是第一条能把一个正常工作的会话移进这种状态的路径。
- 候选修复:在共用的
requireDirectory中加Files.isReadable && Files.isExecutable,+7/−1。由于本 PR 已把requireDirectory移到 resolver 中共用,一行同时覆盖 W2 探针和acquire()认领前的检查,所以也顺带修好了 G0。加上候选补丁后:- 3/3 次变更以
failed/workspace_unavailable失败;后续 Turn 和邻居会话都能完成,下一次变更也能受理。 - G0 指向 000 目录现在 671 ms 内失败(
hosted_turn_failed),不占任何租约。 - 单测在 head 上 486/486、在合并树上 554/554,
checkstyle:check干净。 - 新测试在未打补丁的 head 上失败。
- 补丁见
candidate.patch。
- 3/3 次变更以
T1 — ManagedOperationSchemaUpgradeTest 并未钉住升级修复(测试缺口)。
- 这个测试只读一张空表,
RowMapper根本不会被执行。 - M18 让
additiveString无条件读列;M18b 让additiveLong也无条件读列,这等同于a2d7bd3b02修复之前的行为。这个 H2 测试对两个变异体都照样通过。 - 只有
WorkspaceSessionRetentionMySqlIT能杀死它们。我在 MySQL 8.4.7 上跑过:head 16/16,M18b 下报错。 - 候选补丁中改写后的测试会先插入一行升级前受理的操作,能杀死 M18 和 M18b,在 head 上通过。
次要备注(不影响合并)
- N1「陌生人永远 404」。 在
workspace-files-enabled=false时,陌生人(无任何授权)访问一个真实存在的绑定会话会得到409 workspace_unavailable,而对不存在的 id 得到404,GET也是404。这构成了租户内的存在性探测,因为第 3–4 步排在可见性检查之前。同样配置下,重放一次已完成的变更会得到 409,而不是返回原操作。至于陌生人访问旧式会话得到 400unsupported_feature,这没有问题:旧式会话本来就是租户内可见的(同一个陌生人GET得到 200)。 - N2 回滚。「旧二进制会忽略这些列」对已完成的变更成立:base jar 能在 V34 的库上启动、能读会话,
GET目录变更操作得到400 invalid_request。但回滚时仍处于未结状态的变更,在 base jar 下会一直保持未结,base jar 日志报IllegalArgumentException: No enum constant …OperationKind.CWD_CHANGE;它会阻塞该会话后续的受理,直到 W2 二进制回来,然后由 W2 二进制恰好完成一次。 - N3 JSON 宽松转换。
cwd_relative: 7会被当作"7"受理;expected_context_revision传"1"或1.9都会被当作 1 接受。契约写的是字符串和整数。 - N4。 处于重试退避中的操作(状态 RUNNING、投递 PENDING)读出来是
installing。 - N5 与 feat(managed-agent): let a Workspace-bound Session's creator submit, cancel and rename #13112 的握手(已记录为后续)。 后续 Turn 的受理不会检查未结的目录变更。真实栈上这是安全的:
- 变更未结期间受理了一个 Turn,且 Turn 在提交时仍在运行 → 变更以
session_context_busy失败,Turn 在旧目录完成,revision 不变。 - 如果 Turn 先结束,变更随后正常提交。
- 变更未结期间受理了一个 Turn,且 Turn 在提交时仍在运行 → 变更以
- N6。 M15(摘要按原始写法而非规范形式计算)在全部 485 个单测下存活,只有
HostedPublicWorkspaceIT能杀死它,而 CI 会跑这个 IT。这里只是记录规范化是在哪里被钉住的。
成立的部分
| 场景 | 结果 |
|---|---|
S1:两侧都是 202 → 0.12 s 内完成;规范化重放(child2//、child2/./、./child2)返回同一 id 且 replayed=true;同键不同目标或 revision → 409 idempotency_conflict;过期 revision → 409 context_revision_conflict;跨侧读取操作;数据库行为 COMPLETED/CONFIRMED 并带回执;每次变更在公开与 WebShell 的实时 SSE 上恰好一条 session.context.changed |
24/24 |
S2:401 actor_required;400 的键 / 请求体 / 类型校验;12 种词法形态在读取任何会话事实之前就得到 invalid_cwd;未知 id、跨租户、陌生人 → 404;读者与其他 Workspace 创建者 → 403 session_operation_forbidden(带过期 revision 也是如此,不泄漏状态);撤销 can_create、DRAINING、generation 漂移 → 409 workspace_unavailable;首个 Turn 活跃 → 409 session_context_busy(CAS 与 403 优先);任何拒绝都不写操作行 |
50 ✓ |
S2b:CLOSING/CLOSED/ARCHIVING/ARCHIVED/DELETING → 409 session_state_conflict,DELETED → 404;已完成变更的重放先于状态门返回。状态用 SQL 写入,因为 macOS 栈上不支持绑定会话的 close。 |
14/14 |
S3:缺失、普通文件、FIFO、指向内部 / 外部的软链、中间段软链、大小写别名、NFD 别名、超过 NAME_MAX 的路径段 → failed/workspace_unavailable,binding、revision、事件都不变,也不重试;NFC、中文加空格、多层嵌套、.、同一目录 → 完成,revision +1 |
16/16 |
| S4:10 轮 × 20 个并发受理(公开与 WebShell 交错)→ 每轮恰好一个 202、一次提交、一条事件;20 个同键并发请求(写法混用)→ 只有一行;同键不同内容 → 409 | 3/3 |
S5:注入两次提交失败 → 重试,attempt_count=2,只提交一次;在提交事务内(探针已通过)kill -9 → 期间读取仍是旧值、操作读作 installing,59.6 s 后(租约到期)以认领代数 2 回收,只有一条事件;在认领阶段 kill -9 → kill 后 8 s 完成;在认领与提交之间撤销 can_create / can_read、改为 DRAINING、generation 变化、删除或替换目录 → 失败,binding 保持,之后的变更仍可用 |
11/11¹ |
| S7:base 上路由 404;V33 → V34 升级;升级前创建的会话能变更目录;回滚与再升级见 N2 | 8/8 |
S9/S13,在 head ⊕ main 2b15eac862 上:变更后的下一个后续 Turn 写入 child2/(而不是 child/);变更被拒后写回旧目录;变更为 . 后写到根目录;提交之后目录被替换为指向外部的软链 → 下一个 Turn 以 hosted_turn_failed 失败,外部目录没有任何写入;Turn 与未结变更竞争的情形见 N5 |
6/6 + 1/1 |
单测(H2)/ HostedPublicWorkspaceIT(H2、MySQL 8.4.7)/ checkstyle |
485/485 · 3/3 · 3/3(另有 1 个仅 Linux 的用例被跳过)· 干净 |
| 变异测试:23 个变异体,覆盖受理顺序、CAS、繁忙屏障、创建者 / Registry 校验、提交时的复查、租约守卫、探针规则、事件和 kind 分支,用 PR 的聚焦测试运行 | 19/23 被杀死;存活的 M15/M18/M18b 见上文;M17(只去掉 failCwdChangeOperation 的租约过期条件,owner / generation 仍在守卫)近似等价 |
¹ D1(撤销 can_read)由操作行确认(FAILED / workspace_unavailable)。探针以已被撤权的 actor 轮询,得到 404,这本身就是正确行为。
证据(装置、探针、MySQL 故障 / 挂起触发器、各臂结果、变异台账、候选补丁):wenshao/qwen-code@84e0818 → pr13247/
Bring in #13112 (bound later turns) and friends; besides the textual resolutions (contract header keeps v1.30.0 with the upstream v1.28 entry, README inventory, migration V34 -> V35 after #13090) this adds the two integrations the merge demands: * insertTurnCommand refuses session_context_busy for a bound Session with an open operation - the busy-barrier handshake the W2 design recorded, pinned both directions by a new store test. * The hosted IT gains the end-to-end pivot proof that #13112 unlocks: a bound later Turn submitted after a settled change writes files only in the committed directory.
|
Merged origin/main ( Textual resolutions: contract header stays v1.30.0 and now carries the upstream v1.28 changelog entry; the README capability inventory is unioned; the migration moved V34 → V35 after #13090 took V34. Two integrations beyond text, because #13112's merged admission path did not gate bound later turns on open operations — the handshake this design reserved:
Re-verified on the merged tree (JDK 21, Maven 3.9.11 — the new SpotBugs gate runs in |
… replay Review round 1 on the W2 slice found real leaks before cosmetic items; the critical one was ordering: the idempotency replay and the deployment flag check ran before the creator check, so a revoked grant could replay-read a live operation and an ungranted caller could map Sessions and the flag from the refusal codes. The creator check now runs immediately after the binding gate, matching the sibling lifecycle admission's invisible-caller discipline; the contract descriptions and the design admission table enumerate the wired order truthfully (including the 401 step). * failCwdChangeOperation reports whether the terminal write landed, so a lease-lost refusal is logged as contested instead of as a terminal refusal that never happened; settlement timestamps now anchor on database time like the rest of the lifecycle queue. * The legacy getPublic/getWebShell operation reads are gone; the kind-aware pair is the only projection. * Test coverage closes the gaps the same review raised: probe-per-conjunct failure fences, recovery-scan visibility for cwd rows, a null-binding settlement guard, the added-column upgrade path reading a real row, the WebShell failed projection carrying failureCode, and the busy barrier refusing a lifecycle admission from the other side.
Review round 2 measured four Critical gaps against the merged W2 slice; every one of them changed production order or a guard, not a doc. * The idempotency contract outranks the deployment gate: replay now runs before the workspace-files check, so a lost 202 resolves through the original operation even when execution has since been disabled, while the actor invisibility check still precedes both. * The settlement probe verifies the candidate binding instead of the current one — a change away from a destroyed current directory stays reachable, which is the escape the feature exists for — and the shared directory rule carries the worker's own access(R_OK|X_OK) checks so nothing passes a probe that the next turn would refuse as an untyped wedge. * The bound later-Turn barrier counts context-changing operation rows only, so a stuck display mutation or an in-flight ACTION_RESPONSE can never wedge the Session's only re-acquire path; the wider barrier stays where the operation machinery owns delivery. * The admission service enforces 401-before-key itself, the refusal log carries the broker code, class and message, and the cwd refusals use error.getCode() instead of a hard-coded literal. Tests and documents catch up: pipeline and citations rewritten to the wired truth, the hosted IT runs a real initial tool Turn for the later Turn to escape with fail-fast diagnostics, and the contract texts and generated types name unsupported_feature, the deleted-404 and the over-128 classification divergence where they belong.
R2 round disposition (fix commit
|
… probe's transient class Review round 3 measured one Critical and a set of acceptance-level gaps against the merged W2 slice; each is answered in code, not prose. * A requested permission Action holds the Session busy at both cwd admission and the settlement re-check, the same state the sibling lifecycle refuses with 409 turn_active - committing under an answerable approval would certify a context certainty that does not exist. * A binding detached between the coordinator read and the commit now answers workspace_unavailable, not context_revision_conflict. * Momentary probe I/O failures keep their cause and retry through the delivery machine (unavailableTransient + isRetryable consulted at settlement) instead of certifying a verification that never happened; structural refusals remain the typed terminal workspace_unavailable. Test hardening from the same round: action-blocked admission/settlement with answered-or-expired release, settlement busy pinned on a second open operation (the previously untested disjunct), the two permission conjuncts discriminated mode by mode with a uid-aware gate, the shared rule proven to fence acquire() before installContext or ownership.claim, the guarded probe pinned for the negative half (no current-binding call slips through), key-form refusals probed on both surfaces including the 129-char divergence, schema-shape validation of the cwd operation DTOs against the reviewed OpenAPI schemas including the failed => failure_code conditional, the reclaim tail asserting the settled row is re-read after the second dispatch, and the sentinel that makes the later-Turn pivot individually discriminating. Docs and contract texts match the shipped truth.
R3 round disposition (fix commit
|
Real-stack re-verification (round 2) @
|
| Mutant | Unit | Hosted IT | Real stack |
|---|---|---|---|
| N5: admission ignores a requested Action (the admission half of R3-1) | survives | survives | S14-C admits a 202 |
| N12: a transient probe refusal fails terminally | survives | survives | S15-A ends FAILED in 15 ms |
| N11: a mount-continuity I/O failure is classified structural | survives | survives (rerun) | S15-A ends FAILED |
N10: a requireDirectory I/O failure is classified structural |
survives | survives | reachable only through a race between isDirectory and toRealPath |
| M15: digest over the raw spelling | survives | killed | — |
Why the unit tests miss them:
aRequestedActionBlocksAdmissionAndSettlementplants its Action on a Session that already has an open cwd operation, sohasOpenOperationrefuses before the Action clause is reached.coordinatorSettlesFailsTerminallyAndRetriesTransientErrorsinjects anIllegalStateException, not the new retryableRuntimeBrokerException.- Nothing removes a mount root.
candidate-tests.patch adds three tests (+65 lines, test-only): an Action alone blocks admission; a unavailableTransient refusal keeps the operation open with attempt_count=1; a missing mount root is isRetryable() and heals when the root returns. They pass on head (37/37, checkstyle clean) and kill N5, N12 and N11 respectively.
O2: the transient retry has no attempt bound (decide or document). In S15 the mount root disappears for 9 s mid-settlement:
- While it is missing, the operation retries (
attempt_count3). The Session refuses later Turns and further changes (409 session_context_busy) and renames (409 session_operation_active). By code, close/archive/delete hit the same barrier; I could not exercise that because bound close is unavailable on macOS. - When the same root returns, the operation completes. A replacement root with a new file key ends
FAILED workspace_unavailable. - If the root never returns, the operation retries forever: the delay is capped at 60 s but the attempt count is not. The comment at
SessionLifecycleCoordinator.java:189calls this the "bounded retry". Healing on return is a reasonable choice. I'd still either cap the attempts and end typed, or reword the comment and document that an unreachable mount holds the Session until an operator restores it.
O3, informational. A readable non-creator on a DELETED Session now gets 403 session_operation_forbidden on the cwd route (round 1: 404). The sibling close/archive/delete routes answer exactly the same, with the creator and strangers getting 404 (S18), so this is consistent with them.
O4, not W2. ManagedArtifactReadIntegrationTest.OverlappingReadsRefuseImmediatelyAndTimeoutReleasesPermit (80 ms read timeout) failed in 2 of 11 head ⊕ main runs and 0 of 10 runs on pure main, under host load of 80–140. Interleaved runs passed 6/6 on each arm, the PR does not touch that path, and the second full-suite run passed. I count it as a load flake.
What holds (head ⊕ main unless noted)
| Scenario | Result |
|---|---|
S1: both surfaces, normalized replay, session.context.changed exactly once, live SSE |
24/24 (head 24/24) |
| S2 admission matrix / S2b status gate (the reader-on-DELETED expectation updated, see O3) | 50 + 4 rig expectations / 13/14 |
| S3 settlement on real APFS shapes (the mode-000 directory is now refused) | 15/16 + F1 fix |
| S4: 10 rounds × 20 concurrent admissions / S5: retry, kill -9 inside the commit (reclaimed after 58.7 s, claim generation 2, one event), kill -9 inside the claim, facts moving between claim and commit | 3/3 / 11/11¹ |
| S7: main jar V34 → V35 → rollback → roll-forward | 8/8 |
| S9/S13/S17: the next later Turn writes into the new directory; a symlink swapped in after commit → typed failure; Turn, rename and cancel against an open change | 7/7 · 1/1 · 3/3 |
S14 permission-Action gate: waiting approval → 409; cancel or approve → released; requested Action without a Turn → 409; past expiresAt → admitted |
5/5² |
| S16: a change away from a destroyed current directory completes, and the next Turn writes into the new one | 1/1 |
clean verify checkstyle:check with SpotBugs, and HostedPublicWorkspaceIT on H2 and MySQL 8.4.7 |
head 585/585 · 3/3 · 3/3; head ⊕ main 591 run, 0 failures (1 skipped) · 3/3 · 3/3 |
¹ D1 (can_read revoked) was read from the operation row (FAILED / workspace_unavailable). The probe polls as the revoked actor and, correctly, gets 404.
² Case C forces the Turn row terminal with SQL so that only the requested Action remains. S19 likewise seeds its PENDING rename and ACTION_RESPONSE rows with SQL.
Not covered:
- I ran nothing on Linux durable local-process, which is now main's default. So the storage-guard half of R2-2 (verify the candidate binding) was not exercised on a real stack. It is pinned in unit tests: mutant N9 (verify the current binding instead) is killed by
aGuardedMountVerifiesTheProbeTarget. - The
head ⊕ mainarm runs the main-builtdist/cli.js. The PR itself changes no CLI code, only the generated WebShell types.
Evidence (rig, probes, per-arm results, mutation ledger and logs, candidate tests): wenshao/qwen-code@547aa9c → pr13247/r2/ · Round 1: issuecomment-5964450917
中文版
真实栈复验(第 2 轮)@ 4f9e60b300
结论:第 1 轮的合并事项都已解决,机制仍然成立,没有发现新的阻塞问题。 B1、F1、T1 已修复并在真实栈上复验,N1、N5 也已修复。合并前有三件可选的事:
- 补上候选测试(+65 行)。R3 新增的两条主张没有任何单测或 Hosted IT 钉住,只有真实栈能发现。
- 协调 V35 编号。
- 对没有次数上限的瞬时重试(O2)做个决定。
目前卡合并的是 bot 停留在 259790e12d 上的 CHANGES_REQUESTED。它针对本 head 的评审还在跑,其余 CI 全绿。
环境:
- 对照臂:
head 4f9e60b300head ⊕ main 2c591ecc08:main 前进了 19 个提交,合并无冲突,W2 迁移为 V35。
- macOS 宿主、JDK 21、MySQL 8.4.7、真实 Spring fat jar(内嵌 Broker)、打包的 Hosted Harness、fixture 模型、rig 用的 actor 适配器。
- 所有场景都在
head ⊕ main上重跑了一遍,关键场景也在head上跑了。 - main 现在包含 feat(managed-agent): default durable local-process and trusted reboot recovery on #13211(durable local-process 默认开启,要求 Linux)。在这个 macOS 装置上,
head ⊕ main臂显式设置了durable-local-process=false和trusted-local-reboot-recovery=false。这是装置配置,不是 PR 的问题。
第 1 轮事项
- B1,已解决。
head ⊕ main合并无冲突、能正常启动。main jar 建出的库(V34)能升级到 V35,main jar 创建的会话在升级后可以变更目录(S7 8/8)。开放中的 feat(managed-agent): broker authentication and broker-provisioned writer credentials #13210 和 fix(managed-agent): stop database amplification on session hot paths #13217 现在也占用 V35,后合入的一方需要改号。 - F1,已修复(R2-3)。
- 权限为 000、444、111 的目录现在都以
failed/workspace_unavailable结束(3/3)。下一个后续 Turn 和同一存储上的邻居会话都能完成,下一次变更也能受理。 - 创建(G0)时直接指向 000 目录,现在 1.45 s 内以
hosted_turn_failed失败,不占任何存储租约。第 1 轮时这个 Turn 会一直 RUNNING 并占着租约。
- 权限为 000、444、111 的目录现在都以
- T1,已修复。 升级测试现在会先插入一行。M18b(去掉对新增列的容错)直接被这个 H2 测试杀死。
- N1,已修复(R2-1)。 关闭 opt-in 时,陌生人得到 404,与不存在的 id 一样;重放已完成的变更返回
202 replayed(S8 4/4)。 - N5,已修复(握手)。
- 变更未结期间提交后续 Turn 得到
409 session_context_busy,变更结束后又能正常受理。 - 卡住的 PENDING 改名,或者未结的
ACTION_RESPONSE操作,不会挡住后续 Turn,但会挡住目录变更(R2-4,S19 3/3)。 - 变更期间取消一个已结束的 Turn、或者改名(
409 session_operation_active),行为都符合设计(S17 3/3)。
- 变更未结期间提交后续 Turn 得到
- N2/N3/N4/N6,未变(次要)。 回滚时未结的操作仍保持未结,直到 W2 二进制回来;标量 JSON 转换仍然宽松;退避期间仍读作
installing;M15 仍然只被HostedPublicWorkspaceIT钉住。
本轮新发现
O1:R3 有两条主张只被真实栈钉住(测试缺口)。 我对新旧 W2 守卫做了 32 个变异体,PR 的聚焦测试杀死 27 个。存活的变异体再分别用全部单测(585 个)、HostedPublicWorkspaceIT、以及装了变异 jar 的真实栈来跑:
| 变异体 | 单测 | Hosted IT | 真实栈 |
|---|---|---|---|
| N5:受理时忽略已请求的 Action(R3-1 的受理那一半) | 存活 | 存活 | S14-C 受理返回 202 |
| N12:瞬时的探针拒绝被当成终态失败 | 存活 | 存活 | S15-A 15 ms 内 FAILED |
| N11:挂载连续性检查的 I/O 失败被当成结构性错误 | 存活 | 存活(重跑确认) | S15-A FAILED |
N10:requireDirectory 的 I/O 失败被当成结构性错误 |
存活 | 存活 | 只有 isDirectory 与 toRealPath 之间的竞态能走到 |
| M15:摘要按原始写法计算 | 存活 | 被杀死 | — |
单测没抓到的原因:
aRequestedActionBlocksAdmissionAndSettlement把 Action 放在一个已有未结目录变更的会话上,所以先被hasOpenOperation拒绝,根本走不到 Action 那一条。coordinatorSettlesFailsTerminallyAndRetriesTransientErrors注入的是IllegalStateException,而不是新的可重试RuntimeBrokerException。- 没有任何测试去移走挂载根目录。
candidate-tests.patch 只加测试,共 3 个(+65 行):单独一个 Action 就能挡住受理;unavailableTransient 拒绝会让操作保持未结、attempt_count=1;挂载根缺失时 isRetryable() 为真,根目录回来后恢复。它们在 head 上通过(37/37,checkstyle 干净),并分别杀死 N5、N12、N11。
O2:瞬时重试没有次数上限(需要决定或写进文档)。 S15 中在结算期间让挂载根消失 9 秒:
- 根目录缺失期间,操作在重试(
attempt_count到 3)。会话拒绝后续 Turn 和再次变更(409 session_context_busy),也拒绝改名(409 session_operation_active)。按代码,close/archive/delete 也会被同一道屏障挡住;这一点我没能实测,因为 macOS 栈不支持关闭绑定会话。 - 同一个根目录回来后,操作完成;换成一个新的根目录(文件 key 不同)则以
FAILED workspace_unavailable结束。 - 如果根目录一直不回来,就会永远重试:间隔封顶 60 s,但次数没有上限。
SessionLifecycleCoordinator.java:189的注释把这称为「有界重试」。根目录回来能自愈是合理的选择,但建议二选一:要么给次数设上限并以带类型的错误结束;要么改掉注释的措辞,并在文档里写明挂载不可达时会话会一直被占住,直到运维恢复。
O3,仅供参考。 有读权限的非创建者在已删除的会话上调用目录变更,现在得到 403 session_operation_forbidden(第 1 轮是 404)。同级的 close/archive/delete 路由答复完全一样,创建者和陌生人都是 404(S18),所以这与它们保持一致。
O4,与 W2 无关。 ManagedArtifactReadIntegrationTest.OverlappingReadsRefuseImmediatelyAndTimeoutReleasesPermit(读超时 80 ms)在宿主负载 80–140 下,head ⊕ main 跑了 11 次失败 2 次,纯 main 跑了 10 次失败 0 次。交替运行时两臂各 6/6 通过,PR 也没有改这条路径,第二次全量运行通过。我把它归为负载导致的偶发失败。
成立的部分(除注明外均为 head ⊕ main)
| 场景 | 结果 |
|---|---|
S1:两侧接口、规范化重放、session.context.changed 恰好一条、实时 SSE |
24/24(head 24/24) |
| S2 受理矩阵 / S2b 状态门(读者访问已删除会话的预期已更新,见 O3) | 50 + 4 条装置预期 / 13/14 |
| S3 真实 APFS 上的各种目录形态(000 目录现在会被拒绝) | 15/16 + F1 修复 |
| S4:10 轮 × 20 个并发受理 / S5:重试、在提交事务内 kill -9(58.7 s 后回收,认领代数 2,只有一条事件)、在认领阶段 kill -9、认领与提交之间事实变化 | 3/3 / 11/11¹ |
| S7:main jar V34 → V35 → 回滚 → 再升级 | 8/8 |
| S9/S13/S17:变更后的下一个后续 Turn 写入新目录;提交后目录被换成软链 → 带类型失败;Turn、改名、取消与未结变更的交互 | 7/7 · 1/1 · 3/3 |
S14 权限 Action 门:等待审批时 → 409;取消或批准后 → 放行;没有 Turn 但 Action 仍是 requested → 409;过了 expiresAt → 受理 |
5/5² |
| S16:从已被删除的当前目录变更出去能完成,下一个 Turn 写入新目录 | 1/1 |
clean verify checkstyle:check(含 SpotBugs)、HostedPublicWorkspaceIT(H2、MySQL 8.4.7) |
head 585/585 · 3/3 · 3/3;head ⊕ main 跑 591 个、0 失败(跳过 1 个)· 3/3 · 3/3 |
¹ D1(撤销 can_read)是从操作行读到的(FAILED / workspace_unavailable)。探针以已被撤权的 actor 轮询,得到 404,这本身是正确行为。
² 场景 C 用 SQL 把 Turn 行改成终态,好让只剩下那个 requested 的 Action。S19 里卡住的 PENDING 改名和 ACTION_RESPONSE 行同样是用 SQL 写入的。
未覆盖:
- 我没有在 Linux durable local-process 上跑任何场景,而它现在是 main 的默认配置。所以 R2-2(校验的是候选 binding)中涉及 storage guard 的那一半,没有在真实栈上验证。单测钉住了它:变异体 N9(改为校验当前 binding)会被
aGuardedMountVerifiesTheProbeTarget杀死。 head ⊕ main臂使用的是 main 构建出的dist/cli.js。PR 本身没有改 CLI 代码,只改了生成的 WebShell 类型。
证据(装置、探针、各臂结果、变异台账与日志、候选测试):wenshao/qwen-code@547aa9c → pr13247/r2/ · 第 1 轮报告:issuecomment-5964450917
V35 was claimed by V35__managed_session_tool_profile on main while this branch was under review; this keeps `V36__managed_cwd_operation.sql` as the version Flyway can place without colliding, together with the comments that cite it and the merged-origin shape already verified green.
|
Merged origin/main ( Re-verified on the merged tree (fenced |
Real-stack re-verification (round 3) @
|
| Scenario | Result |
|---|---|
| S1 both surfaces / S4 races / S9, S13, S17 later Turns and interplay | 24/24 · 3/3 · 7/7 · 1/1 · 3/3 |
| S2 / S2b / S3. Same expectation updates as round 2: the tenant filter answers 400, bound close is unavailable on macOS, a reader on DELETED gets 403 like the sibling routes, and the 000 directory is now refused. | 50 · 13/14 · 15/16 (+ the expected ones) |
| F1 stays fixed: 000/444/111 refused while the later Turn and the neighbour complete; G0 into a 000 directory fails typed in 733 ms with 0 leases held | 3/3 |
| S5: retry; kill -9 inside the commit (reclaimed after 59.4 s, claim generation 2, one event); kill -9 inside the claim; facts moving between claim and commit | 11/11¹ |
S7: main jar V39 → head ⊕ main V40 → rollback → roll-forward |
8/8 |
| S8 opt-in off / S14 permission-Action gate / S15 transient mount root / S16 destroyed current directory / S19 barrier population | 4/4 · 5/5 · 2/2 · 1/1 · 3/3 |
| S20 (new): same-tenant admission bursts, the deadlock shape #13365 fixed. 4 rounds of 8 cwd changes + 8 later Turns + 8 creations, released together. | 96/96 got 202, every change completed once; InnoDB lock_deadlocks 0 → 0 |
clean verify checkstyle:check with SpotBugs; HostedPublicWorkspaceIT on H2 and MySQL 8.4.7 |
head: 607 run, 0 failures · 3/3 · 3/3. merge: 676 run, 0 failures · 3/3 · 3/3² |
¹ D1 (can_read revoked) was read from the operation row (FAILED / workspace_unavailable). The probe polls as the revoked actor and, correctly, gets 404.
² One H2 run errored in ownerAnswersHostedApprovalsThroughBothSurfaces, an approval test rather than a W2 one: the Turn was still RUNNING after 35 s at host load ~45. Three reruns passed on the merge, and three on main as a control; the MySQL run passed.
Still open from round 2 (non-blocking):
- O1: mutants N5 (admission ignores a requested Action), N11 and N12 (transient I/O classified structural, or a transient refusal made terminal) still survive the PR's tests on this head. The round-2
candidate-tests.patchapplies cleanly, passes 37/37 with checkstyle clean, and kills all three. - O2: the transient retry still has no attempt bound. S15 behaves exactly as in round 2.
Setup:
- macOS host, JDK 21, MySQL 8.4.7, the real Spring fat jar with its embedded Broker, the packaged Hosted Harness built from the merge tree, and a fixture model.
- The merge arm sets
durable-local-process=falseandtrusted-local-reboot-recovery=false, as main's feat(managed-agent): default durable local-process and trusted reboot recovery on #13211 requires on non-Linux hosts. - Not covered: Linux durable local-process.
Evidence: wenshao/qwen-code@c7e1cf9 → pr13247/r3/ · Round 2: issuecomment-5976428811 · Round 1: issuecomment-5964450917
中文版
真实栈复验(第 3 轮)@ 1a55a6b246
结论:又出现阻塞,迁移编号再次与 main 撞号(B1)。
- fix(managed-agent): stop database amplification on session hot paths #13217 在 13:57 UTC 合入,带有
V36__managed_session_journal_activation.sql(以及 V37–V39),比本 PR 改号为 V36 晚 40 分钟。PR 自己的「Flyway migration version uniqueness」检查在 13:20:49 UTC 就已跑完,所以至今仍显示绿色。 - GitHub 显示没有文本冲突。如果直接合并,main 上会出现两个 V36 文件:服务启动失败,main 推送时的唯一性检查也会变红。
- 把 W2 迁移改号之后,
head ⊕ main上其余各项全部成立。 - 第 2 轮的可选事项(候选测试、没有次数上限的瞬时重试)仍未处理;W2 的源码和测试自
4f9e60b300起没有改动。
B1,在 head ⊕ main 17c182eda0 上(比 PR 的合并基多 11 个提交):
- 真实 fat jar 启动即失败,报
Found more than one migration with version 36(V36__managed_session_journal_activation.sql与V36__managed_cwd_operation.sql)。全新的 MySQL 库、以及 main jar 建出的库(V39)都是如此。 - 在同一棵树上,CI 自带的
scripts/check-flyway-migrations.js返回 1;只看 head 时它是通过的。 - 把 W2 文件改名为 V40 后服务能启动。唯一性脚本在
17c182eda0和更新的 main292c49ec4a(没有新增迁移)上都通过;在后者上契约测试与 W2 测试 40/40 通过。 - V40 也不安全: 开放中的 feat(managed-agent): add reliable ACTIVE Workspace deletion (L3) #13354、feat(managed-agent): H3 background Shell and Monitor runtime #13265、feat(managed-agent): broker authentication and broker-provisioned writer credentials #13210、feat(runtime): add experimental Kubernetes CSI runtime and durable worker ACK #13289 都占用了 V40(V41–V43 也已被占)。
- 建议:在合并前一刻改成下一个空闲编号,再针对当时的 merge ref 重跑唯一性检查或 CI。main 最新一次推送之前的绿色检查,说明不了合并后的结果。
其余各项,在 head ⊕ main(V40)上重跑:
| 场景 | 结果 |
|---|---|
| S1 两侧接口 / S4 并发竞争 / S9、S13、S17 后续 Turn 及交互 | 24/24 · 3/3 · 7/7 · 1/1 · 3/3 |
| S2 / S2b / S3。与第 2 轮相同的预期调整:租户过滤器返回 400、macOS 上不支持关闭绑定会话、读者访问已删除会话得到 403(与同级路由一致)、000 目录现在会被拒绝。 | 50 · 13/14 · 15/16(另加上述预期内的项) |
| F1 仍然修复:000/444/111 被拒,后续 Turn 与邻居会话都能完成;G0 指向 000 目录在 733 ms 内以带类型的错误失败,不占租约 | 3/3 |
| S5:重试;在提交事务内 kill -9(59.4 s 后回收,认领代数 2,只有一条事件);在认领阶段 kill -9;认领与提交之间事实变化 | 11/11¹ |
S7:main jar V39 → head ⊕ main V40 → 回滚 → 再升级 |
8/8 |
| S8 关闭 opt-in / S14 权限 Action 门 / S15 挂载根瞬时缺失 / S16 当前目录已被删除 / S19 屏障覆盖范围 | 4/4 · 5/5 · 2/2 · 1/1 · 3/3 |
| S20(新增): 同租户受理突发,即 #13365 修过的死锁形态。4 轮,每轮 8 个目录变更 + 8 个后续 Turn + 8 个会话创建同时放出。 | 96/96 返回 202,每次变更都恰好完成一次;InnoDB lock_deadlocks 0 → 0 |
clean verify checkstyle:check(含 SpotBugs);HostedPublicWorkspaceIT(H2、MySQL 8.4.7) |
head:跑 607 个、0 失败 · 3/3 · 3/3。合并树:跑 676 个、0 失败 · 3/3 · 3/3² |
¹ D1(撤销 can_read)是从操作行读到的(FAILED / workspace_unavailable)。探针以已被撤权的 actor 轮询,得到 404,这本身是正确行为。
² H2 有一次运行在 ownerAnswersHostedApprovalsThroughBothSurfaces 报错,这是审批测试而不是 W2 测试:宿主负载约 45 时,Turn 过了 35 s 仍是 RUNNING。合并树重跑 3 次、main 对照重跑 3 次全部通过;MySQL 那次运行也通过。
第 2 轮遗留(不阻塞):
- O1: 变异体 N5(受理时忽略已请求的 Action)、N11 和 N12(瞬时 I/O 被当成结构性错误、瞬时拒绝被当成终态)在本 head 上仍然能通过 PR 自己的测试。第 2 轮的
candidate-tests.patch可以干净地打上,37/37 通过、checkstyle 干净,并能杀死这三个变异体。 - O2: 瞬时重试仍然没有次数上限,S15 的表现与第 2 轮完全一致。
环境:
- macOS 宿主、JDK 21、MySQL 8.4.7、真实 Spring fat jar(内嵌 Broker)、从合并树构建的打包 Hosted Harness、fixture 模型。
- 合并臂按 main 的 feat(managed-agent): default durable local-process and trusted reboot recovery on #13211 对非 Linux 宿主的要求,设置了
durable-local-process=false和trusted-local-reboot-recovery=false。 - 未覆盖:Linux durable local-process。
证据:wenshao/qwen-code@c7e1cf9 → pr13247/r3/ · 第 2 轮:issuecomment-5976428811 · 第 1 轮:issuecomment-5964450917
|
复核提交: 结论:当前仍有 1 个合并阻断;其余审阅路径未发现新增阻断问题。 [P1 / Critical] Flyway V36 撞号仍然成立。 本 PR 的 这是对已有 B1 报告的独立确认,不重复开新问题。 此前其余 6 条 Critical 已逐项对照当前源码复核:撤权后的重放保持不可见;重放先于部署开关;目录探针校验候选 binding;可读/可搜索检查先于存储 claim;后续 Turn 屏障排除展示 mutation 和 ACTION_RESPONSE;cwd 准入与结算均检查待决 Action。另检查了 revision CAS、租约与 claim generation、失败时保留原目录、提交与事件的事务一致性、两侧 operation DTO 和双语设计说明。 本轮独立验证:JDK 21 / Maven 3.9.11,使用隔离 Maven 仓库构建当前提交的依赖;7 个相关测试类共 70 个仓库测试通过,另 3 个独立探针通过(撤权后不同 payload 的重放仍为 404、关闭开关后重放已完成 operation、仅有待决 Action 时拒绝 cwd 准入)。 验证边界:本轮没有重跑完整 Hosted/MySQL 栈;真实 verified-recovery storage guard 的额外探针被 macOS 平台限制拒绝(要求 Linux),候选 binding 路径仅有源码检查及现有 mock 测试证据,不声称完成了 Linux 真实守卫验证。 已报告的测试覆盖缺口和无限瞬时重试取舍继续留在原讨论中,本轮不把它们重复列为新发现。 |
…rigin Review round 4 measured the transient/terminal split at origin type and found it backfiring the other way: a permanently gone mount root is a NoSuchFileException - an IOException - which mapped to the retryable arm and wedged the Session with unbounded retried scans. Classification is now by the meaning: a vanished mount root or target directory is the structural verdict the terminal `workspace_unavailable` exists to name (NoSuchFileException → unavailable()), while only genuine I/O blips retry (ESTALE/EIO, `unavailableTransient`). * A requested permission Action now has its own admission witness - 409 session_context_busy asserted with no open operation and no active Turn, so the hasDecidableAction conjunct cannot hide behind the || short-circuit. * The settlement's status re-check is pinned: a Session that closed between claim and commit fails typed with the revision and directory untouched. * The sealed-acquire test now also proves no claim is recorded after the refusal (`holder_key IS NOT NULL` count 0), so the ordering the name claims is tested. * The published key-form sentence counts both divergences (over-length and blank) and four more probes pin them on both surfaces; the cwd instance suite also validates the cwd operations against the serving oneOf unions, and its Javadoc stops claiming exclusivity. * The coordinator's retry prose says unbounded-with-capped-delay now, matching the delivery machine.
Main landed the session-journal migration chain meanwhile with its own V36-V39 additions, so the earlier V36 number collided; V40 sits past the new top without touching the journal sequence.
The round-4 fixes put the machinery in but under-pinned it in two places, spotted during this round's self-audit: * The resolver's terminal refusals now assert the retryable flag, and a new case deletes the target and then the mount root to prove a NoSuchFileException verdict is terminal, not retried (the Critical's missing regression witness). * The delivery machine's isRetryable() consult gets its own witness: a retryable probe refusal leaves the operation RUNNING with no failure code and re-settles completed after the backoff, so dropping the consult turns the suite red. * The instance suite's Javadoc no longer claims exclusivity over the record/schema drift harness - the javadoc reword the round-4 commit message announced but never included.
|
Review round 4 disposition (1 Critical + 8 Suggestions — all 9 answered, threaded replies posted and resolved):
Verification on |
Conflicts resolved: - managed-agent-public-api.openapi.json: #13210 took v1.30.0 with its gateway-free qwenSignature scheme this morning; the W2 cwd contract now ships v1.31.0, with both changelog sentences kept and the contract's two (v1.30) feature references re-anchored to (v1.31). - README.md: keep both feature sections - main's broker authentication and writer credentials first, the W2 controlled cwd change after, with the settle retry budget recorded in its refusal paragraph. - web-shell generated client: regenerated from the resolved contract, so it carries the v1.31 text and main's 413 payload_too_large arm. The cwd migration is V45 after upstream took V40-V44 this morning.
The round-6 re-report was right on both surfaces of the third narrowing. ENOTDIR and ELOOP surface as a bare FileSystemException on some JDKs, so the shared rule's residual IOException arm still classified them as retryable; and because the shared rule had also become the throwing-call shape on the acquire path, a persisted plain-file/sub cwd answered retryable there too - the regression this diff tags - turning an immediate typed failure into transientFailure retries or a RECOVERY_BLOCKED park. Split along the duality the guard already has: the acquire path keeps its long-standing boolean-predicate requireDirectory and verifyMountIntact with every I/O anomaly terminal, while the probe twins requireDirectoryForProbe/verifyMountIntactForProbe do the classification by verdict - the residual IOException arm walks the ancestor chain (hasStructuralAncestor): a plain-file or a dangling/looping symlink ancestor is structural, whatever remains (ESTALE/EIO, ENAMETOOLONG) retries through the 8-attempt budget. Witnesses per the round's acceptance, both mutation-proven: assertProbeRefused(plain-file/lib) and (loop-a/lib) beside the nested suite (terminal code and !isRetryable), and a WorkspaceRuntimeTest acquire case on a plain-file/sub session (assertUnavailable's !isRetryable() tripwire plus rival claim). Fix-removed mutant puts both red; restored, both green.
|
Review round 6 disposition (1 inline Critical, third narrowing of R4-1 — answered, threaded reply posted and resolved; the round's deferred list was recorded without requests):
Verification on |
The step has run at the 12-minute edge for weeks - the previous green run on 12d6c41 finished in 11m04s with 684 module tests - and the upstream merge wave (CSI chain, journal chain, H0c) lifted the suite to 990 tests plus the hosted-mysql failsafes, so the step now times out at 12 minutes while maven completes BUILD SUCCESS five minutes later, killing the lane with everything green. Bubble the production workload, not the test outcome.
Real-stack re-verification (round 4) @
|
| Scenario | Result |
|---|---|
| S1 both surfaces / S4 races / S9, S13, S17 later Turns and interplay / S14 Action gate / S8 opt-in off / S19 barrier population | 24/24 · 3/3 · 7/7, 1/1, 3/3 · 5/5 · 4/4 · 3/3 |
| S2 / S2b (same expectation updates as earlier rounds) | 50 · 13/14 |
S3 probe verdicts, including the new structural shapes: a file ancestor (ENOTDIR), a symlink-loop ancestor (ELOOP), a mode-000 parent (EACCES) and a dangling-symlink ancestor all end failed/workspace_unavailable with attempt_count 0 |
19/19 |
| F1 stays fixed: change to a 000/444/111 directory refused, while the later Turn and the neighbour complete. G0 into a 000 directory fails typed in 671 ms with 0 leases held. | 3/3 |
| S15b: mount root vanishes mid-settlement → terminal at once (attempt_count 0); once the root is back, a re-issued change completes | 2/2 |
S22 acquire path (R6 fix): a committed cwd whose ancestor later becomes a regular file or a dangling symlink makes the next later Turn fail hosted_turn_failed in 471 / 729 ms, with 0 leases held; changing back to child recovers |
4/4 |
| S5: retry, kill -9 inside the commit (reclaimed after 59.1 s, claim generation 2, one event), kill -9 inside the claim, facts moving between claim and commit | 11/11¹ |
| S7: main jar V44 → head V45 → rollback → roll-forward | 8/8 |
| S20 same-tenant bursts (the #13365 shape): 4 rounds of 8 cwd changes + 8 later Turns + 8 creations | 96/96 got 202; lock_deadlocks 0 → 0 |
clean verify checkstyle:check with SpotBugs; HostedPublicWorkspaceIT on H2 and MySQL |
990 run, 0 failures (1 skipped) · 3/3 · 3/3 |
| Mutation: 15 mutants of the R4–R6 code plus the earlier survivors, against the focused tests | 12/15. The round-2 survivors are now killed: admission ignoring a requested Action, missing mount root becoming retryable, retryable refusal becoming terminal. The budget, its off-by-one, the vanished-target/EACCES/ancestor arms and the probe's READ/EXECUTE checks are all pinned. |
¹ D1 (can_read revoked) was read from the operation row (FAILED / workspace_unavailable). The probe polls as the revoked actor and, correctly, gets 404.
F2: a deterministic input error is treated as a momentary fault (minor)
The admission's lexical rule allows 1024 code points per path but sets no per-component length limit. A component of 256 or more ASCII characters therefore passes admission and makes readAttributes throw a plain FileSystemException ("File name too long"). The probe files that under "momentary" and retries it through the budget.
S21 on a real stack, cwd_relative = "a" × 300:
- Probes at t = 0, 1, 2, 4, 8, 17, 32, 65 s; it ends
FAILED workspace_unavailableat 124 s. - Throughout, the creator's later Turn and any other change get
409 session_context_busy, and every deferral logs a full stack trace. - Afterwards the Session is unchanged and usable.
This is what aMomentaryIoFailureClassifiesRetryable pins on purpose: it uses "a".repeat(256) as the only deterministic way to reach the transient arm. The input is permanent, though, and any creator can reach it through the API. On APFS, an 86-character CJK name (258 bytes) gets NoSuchFileException and is terminal in 1 s. On Linux ext4 (not run here), a component of 256 or more bytes is expected to hit ENAMETOOLONG.
Option: cap each component at 255 UTF-8 bytes in the lexical rule, giving 400 invalid_cwd at admission and writing no operation row. The transient-arm witness would then need an injected IOException instead.
T3: the acquire-side half of the F1 fix has no unit test (test gap)
Mutants R6 (drop isReadable from the acquire path's requireDirectory) and R7 (drop isExecutable) survive all 990 unit tests and the Hosted IT. The reason: refusesAnUnreadableSessionDirectoryBeforeClaimingStorage only uses mode 000, which fails both checks at once. On the real stack (S23, G0 creation straight into the directory):
| Arm | 111 (search, no read) | 444 (read, no search) | 000 |
|---|---|---|---|
| head | FAILED hosted_turn_failed, 1.2 s, 0 leases |
FAILED, 0.7 s, 0 leases | FAILED, 0.7 s, 0 leases |
| R6 jar | RUNNING after 30 s, lease held (the round-1 wedge) | FAILED | FAILED |
| R7 jar | FAILED | RUNNING after 30 s, lease held | FAILED |
candidate-acquire-modes.patch (test-only, +29/−22) runs that same test over 000, 111 and 444, releasing the rival claim each time. It passes on head (26/26, checkstyle clean) and kills R6 and R7.
R8 (the acquire path's catch made retryable) also survives. It is nearly unreachable, because the predicates never throw on ENOTDIR. So the comment above refusesAPlainFileDescendantCwdTerminallyBeforeClaimingStorage overstates what it guards when it says folding the anomaly into the transient arm "reddens this test".
Not covered: Linux durable local-process (main's default), and with it the real verified-recovery storage guard on the probe's candidate binding.
Evidence: wenshao/qwen-code@04e8404 → pr13247/r4/ · Earlier rounds: R3 · R2 · R1
中文版
真实栈复验(第 4 轮)@ 98e1ab90ae
结论:没有阻塞问题。 迁移为 V45,在 main 69d5db2ff2 上唯一。PR 上次合并后 main 没有再前进,所以 head 本身就是合并结果。第 2、3 轮的事项都已关闭:瞬时重试现在有上限,之前的测试缺口(N5/N11/N12)都已被杀死。
两条不阻塞的发现,均有真实栈证据:
- F2: ENAMETOOLONG 被归为可重试,词法合法的超长名字会让创建者的会话被占 124 s。
- T3: 获取路径上的访问检查(防止 F1 那种卡死出现在 Turn 路径上)只有真实栈能钉住。附了候选测试。
合并时的注意事项:开放中的 #13260、#13354、#13265 也占用 V45。后合入的一方必须改号,因此要在最终的 merge ref 上重新检查唯一性。
成立的部分(head 即合并结果;macOS 宿主、JDK 21、MySQL 8.4.7、真实 Spring + 内嵌 Broker + 打包 Harness)
| 场景 | 结果 |
|---|---|
| S1 两侧接口 / S4 并发竞争 / S9、S13、S17 后续 Turn 及交互 / S14 Action 门 / S8 关闭 opt-in / S19 屏障覆盖范围 | 24/24 · 3/3 · 7/7、1/1、3/3 · 5/5 · 4/4 · 3/3 |
| S2 / S2b(与前几轮相同的预期调整) | 50 · 13/14 |
S3 探针判定,含新的结构性形态:祖先是普通文件(ENOTDIR)、祖先是软链环(ELOOP)、父目录权限 000(EACCES)、祖先是悬空软链,都以 failed/workspace_unavailable 结束,且 attempt_count 为 0 |
19/19 |
| F1 仍然修复:变更到 000/444/111 目录被拒,后续 Turn 与邻居会话都能完成。G0 指向 000 目录在 671 ms 内以带类型的错误失败,不占租约。 | 3/3 |
| S15b:结算期间挂载根消失 → 立即终态(attempt_count 为 0);根目录回来后,重新发起的变更能完成 | 2/2 |
S22 获取路径(R6 修复):已提交目录的祖先后来变成普通文件或悬空软链,下一个后续 Turn 在 471 / 729 ms 内以 hosted_turn_failed 失败,不占租约;改回 child 后恢复 |
4/4 |
| S5:重试、在提交事务内 kill -9(59.1 s 后回收,认领代数 2,只有一条事件)、在认领阶段 kill -9、认领与提交之间事实变化 | 11/11¹ |
| S7:main jar V44 → head V45 → 回滚 → 再升级 | 8/8 |
| S20 同租户突发(即 #13365 修过的形态):4 轮,每轮 8 个目录变更 + 8 个后续 Turn + 8 个会话创建 | 96/96 返回 202;lock_deadlocks 0 → 0 |
clean verify checkstyle:check(含 SpotBugs);HostedPublicWorkspaceIT(H2、MySQL) |
跑 990 个、0 失败(跳过 1 个)· 3/3 · 3/3 |
| 变异测试:对 R4–R6 的代码及之前的存活者做了 15 个变异体,用聚焦测试运行 | 12/15。第 2 轮的存活者现在都被杀死:受理时忽略已请求的 Action、挂载根缺失变成可重试、可重试拒绝变成终态。重试次数上限及其差一错误、目标消失 / EACCES / 祖先判定的各个分支、探针的 READ/EXECUTE 检查都已被钉住。 |
¹ D1(撤销 can_read)是从操作行读到的(FAILED / workspace_unavailable)。探针以已被撤权的 actor 轮询,得到 404,这本身是正确行为。
F2:确定性的输入错误被当成瞬时故障(次要)
受理时的词法规则限制整条路径最多 1024 个码点,但没有对单个路径段设长度上限。因此一个 256 个及以上 ASCII 字符的路径段能通过受理,然后让 readAttributes 抛出普通的 FileSystemException(「File name too long」)。探针把它归为「瞬时」,按重试次数上限反复重试。
真实栈上的 S21,cwd_relative = "a" × 300:
- 在 t = 0、1、2、4、8、17、32、65 s 各探测一次,124 s 时以
FAILED workspace_unavailable结束。 - 整个过程中,创建者的后续 Turn 和任何其他变更都得到
409 session_context_busy,每次推迟都会打出一份完整的堆栈。 - 结束后会话不变、可以正常使用。
这正是 aMomentaryIoFailureClassifiesRetryable 有意钉住的行为:它用 "a".repeat(256) 作为触达瞬时分支的唯一确定性手段。但这个输入是永久性的,任何创建者都能通过 API 触发。在 APFS 上,86 个中文字符(258 字节)得到的是 NoSuchFileException,1 s 内就终态。在 Linux ext4 上(本轮未跑),256 字节及以上的路径段预计会得到 ENAMETOOLONG。
可选做法:在词法规则里把每个路径段限制在 255 个 UTF-8 字节以内,受理时直接返回 400 invalid_cwd,不写操作行。瞬时分支的测试则改用注入的 IOException。
T3:F1 修复中获取路径那一半没有单测(测试缺口)
变异体 R6(从获取路径的 requireDirectory 中去掉 isReadable)和 R7(去掉 isExecutable)在全部 990 个单测和 Hosted IT 下都存活。原因是 refusesAnUnreadableSessionDirectoryBeforeClaimingStorage 只用了 000 权限,它同时让两项检查失败。真实栈上(S23,G0 创建时直接指向该目录):
| 臂 | 111(可搜索、不可读) | 444(可读、不可搜索) | 000 |
|---|---|---|---|
| head | FAILED hosted_turn_failed,1.2 s,不占租约 |
FAILED,0.7 s,不占租约 | FAILED,0.7 s,不占租约 |
| R6 jar | 30 s 后仍 RUNNING,占着租约(即第 1 轮的卡死) | FAILED | FAILED |
| R7 jar | FAILED | 30 s 后仍 RUNNING,占着租约 | FAILED |
candidate-acquire-modes.patch(只改测试,+29/−22)让同一个测试分别跑 000、111、444,每次都释放对手的认领。它在 head 上通过(26/26,checkstyle 干净),并能杀死 R6 和 R7。
R8(把获取路径的 catch 改成可重试)也存活。它几乎走不到,因为这些谓词在 ENOTDIR 时不会抛异常。所以 refusesAPlainFileDescendantCwdTerminallyBeforeClaimingStorage 上方的注释说「把这个异常并入瞬时分支会让该测试变红」,是夸大了它实际守住的范围。
未覆盖: Linux durable local-process(main 的默认配置),因此也没有在真实环境下验证 verified-recovery storage guard 对探针候选 binding 的检查。
证据:wenshao/qwen-code@04e8404 → pr13247/r4/ · 之前各轮:R3 · R2 · R1
The merged task-journal chain (#13265) took V45 meanwhile, so the renumber hops again; V46 sits past the new top.
Conflicts resolved: - managed-agent-public-api.openapi.json: the task-journal chain (#13265) took v1.31.0 with its Stage H3 section; the W2 cwd contract now ships v1.32.0 (all three changelog sentences kept: qwenSignature v1.30, task events v1.31, cwd change v1.32), and the two (v1.31) feature references are re-anchored to (v1.32). - web-shell generated client regenerated from the resolved contract. The cwd migration is V46 after the task journal took V45.
|
Re-verified One new blocker — [P1/Critical] Catch host-invalid target paths before they escape into unbounded delivery retries. In WorkspaceRuntimeResolver.java:177–180, On a Windows local-process deployment with workspace files enabled and the Linux-only durable/reboot-recovery options disabled, a creator can submit Evidence on this exact head:
Please include path construction/resolution in the probe's terminal exception handling, so an invalid host path settles as Validation: 84 repository tests passed (0 failures/errors/skips), plus 3 independent probes covering guard I/O classification, candidate validation after deletion of the old cwd, and the retry defect above. The last probe deliberately asserts the defect; its green result is reproduction evidence. Current completed CI checks are green; web-shell E2E Smoke is still running. This was macOS/JDK 21 verification with an independently executed Windows parser and fault injection, not native Windows E2E; Linux native storage identity and a live hosted worker were not exercised locally. The already documented ENAMETOOLONG and acquire-access test suggestions are not new findings here. 中文结论:此前的迁移冲突及已报告的错误分类问题已修复;本轮发现一个新阻塞项——Windows 非法路径可绕过异常捕获和 8 次重试上限,使会话持续被未完成操作锁住。证据由官方路径解析器和精确提交上的故障注入测试组成,未声称运行过 Windows 全栈。 |
Real-stack re-verification (round 5) @
|
Scenario (head ⊕ main) |
Result |
|---|---|
| S1 both surfaces / S4 races / S9, S13, S17 later Turns and interplay / S14 Action gate / S8 opt-in off / S19 barrier population | 24/24 · 3/3 · 7/7, 1/1, 3/3 · 5/5 · 4/4 · 3/3 |
| S2 / S2b (same expectation updates as earlier rounds) / S3 probe verdicts, including the structural ancestors (ENOTDIR, ELOOP, EACCES parent, dangling link) | 50 · 13/14 · 19/19 |
| F1 stays fixed (S11 3/3); G0 into a 111/444/000 directory fails typed in ~0.7 s with 0 leases (S23 3/3); S22 acquire path keeps its terminal verdicts (4/4); S15b; S16 | all pass |
| S5: retry, kill -9 inside the commit and inside the claim, facts moving between claim and commit / S7: main jar V45 → V46 → rollback → roll-forward | 11/11¹ · 8/8 |
| S20 same-tenant bursts (4 rounds of 8 cwd changes + 8 later Turns + 8 creations) | 96/96 got 202; lock_deadlocks 0 → 0 |
clean verify checkstyle:check with SpotBugs; HostedPublicWorkspaceIT on H2 and MySQL |
head: 1009 run, 0 failures · 3/3 · 3/3. merge: 1020 run, 0 failures² · 3/3 · 3/3 |
¹ D1 (can_read revoked) was read from the operation row (FAILED / workspace_unavailable).
² The first merge run had 4 errors, all in Issue13180HardenedVerificationTest (from #13210, not W2): its Tomcat connector failed to bind port 59570, so the Spring context never loaded. Three reruns of that class pass on the merge, three pass on main, and a second full run passed.
On the 05:06 P1 (Windows InvalidPathException): the code shape it describes is real on this head. requireDirectoryForProbe builds Path.of(root) and base.resolve(cwdRelative) outside its try (WorkspaceRuntimeResolver.java:178–179), while the acquire-path twin requireDirectory builds them inside a try that catches IllegalArgumentException. So only the settlement probe can let the exception escape to the coordinator's generic retry. On macOS and Linux the only input that makes Path.resolve throw is NUL, and the lexical rule already rejects control characters with 400 invalid_cwd (S2). This rig cannot reach the case, and I did not run Windows, so I have nothing to add beyond confirming the code shape.
Still open from round 4 (non-blocking), unchanged:
- F2: a lexically valid
"a" × 300component still holds the creator's Session busy for 124 s. Probes run at t = 0, 1, 2, 4, 8, 15, 32, 64 s and it endsFAILED workspace_unavailable(S21). Later Turns and other changes get409 session_context_busymeanwhile. - T3: the acquire path's
isReadable/isExecutablechecks are still pinned only by the real stack. The round-4candidate-acquire-modes.patchstill applies cleanly.
Merge-time caveat: open #13260 and #13354 also claim V46. Re-check uniqueness on the final merge ref.
Not testable yet: #13265's child_run/monitor_run domains are still off in production config. Once they are on, a background Shell outlives its Turn, but the W2 barrier only counts active Turns, open operations and requested Actions. That handshake, and the "fresh Runtime Session per Turn" premise the design relies on, deserve a re-check when H3 is enabled. This note comes from reading the code; I have not tested it.
Evidence: wenshao/qwen-code@3b0e111 → pr13247/r5/ · Round 4: issuecomment-6007931423
中文版
真实栈复验(第 5 轮)@ f583ec3774
结论:真实栈上没有新的阻塞问题,与第 4 轮结论一致。 05:06 发出的 仅限 Windows 的 P1 不在本装置的覆盖范围内(见表格下方说明)。第 4 轮之后 PR 只合并了 main:
- feat(managed-agent): H3 background Shell and Monitor runtime #13265(H3 后台 Shell 与 Monitor)占用了 V45,所以 W2 迁移改为 V46,契约从 1.31.0 升到 1.32.0。
- W2 的 resolver、coordinator、service、词法规则和测试都没有改动,只有升级测试里一处迁移编号注释变了。
main 又前进了 6 个提交,其中包括 #13291(本地 Runtime 工具结果持久化)。所以全部场景都在 head ⊕ main f765bcf5d3 上重跑:合并无冲突,唯一性脚本通过,main 仍为 V45 / 契约 1.31.0。
场景(head ⊕ main) |
结果 |
|---|---|
| S1 两侧接口 / S4 并发竞争 / S9、S13、S17 后续 Turn 及交互 / S14 Action 门 / S8 关闭 opt-in / S19 屏障覆盖范围 | 24/24 · 3/3 · 7/7、1/1、3/3 · 5/5 · 4/4 · 3/3 |
| S2 / S2b(与前几轮相同的预期调整)/ S3 探针判定,含结构性祖先(ENOTDIR、ELOOP、父目录 EACCES、悬空软链) | 50 · 13/14 · 19/19 |
| F1 仍然修复(S11 3/3);G0 指向 111/444/000 目录约 0.7 s 内以带类型的错误失败、不占租约(S23 3/3);S22 获取路径保持终态判定(4/4);S15b;S16 | 全部通过 |
| S5:重试、在提交事务内和认领阶段 kill -9、认领与提交之间事实变化 / S7:main jar V45 → V46 → 回滚 → 再升级 | 11/11¹ · 8/8 |
| S20 同租户突发(4 轮,每轮 8 个目录变更 + 8 个后续 Turn + 8 个会话创建) | 96/96 返回 202;lock_deadlocks 0 → 0 |
clean verify checkstyle:check(含 SpotBugs);HostedPublicWorkspaceIT(H2、MySQL) |
head:跑 1009 个、0 失败 · 3/3 · 3/3。合并树:跑 1020 个、0 失败² · 3/3 · 3/3 |
¹ D1(撤销 can_read)是从操作行读到的(FAILED / workspace_unavailable)。
² 合并树第一次运行有 4 个错误,全部在 Issue13180HardenedVerificationTest(来自 #13210,与 W2 无关):它的 Tomcat connector 绑定端口 59570 失败,Spring 上下文没能加载。该测试类在合并树上重跑 3 次、在 main 上重跑 3 次全部通过,第二次全量运行也通过。
关于 05:06 的 P1(Windows InvalidPathException): 它描述的代码形态在本 head 上确实存在。requireDirectoryForProbe 在 try 之外构造 Path.of(root) 和 base.resolve(cwdRelative)(WorkspaceRuntimeResolver.java:178–179),而获取路径上的 requireDirectory 是在一个会捕获 IllegalArgumentException 的 try 里构造的。所以只有结算探针会让这个异常逃到 coordinator 的通用重试里。在 macOS 和 Linux 上,唯一能让 Path.resolve 抛异常的输入是 NUL,而词法规则已经用 400 invalid_cwd 拒绝了控制字符(S2)。本装置触达不到这个情形,我也没有在 Windows 上跑,因此除了确认代码形态之外没有更多可补充的。
第 4 轮遗留(不阻塞),未变:
- F2: 词法合法的
"a" × 300路径段仍会让创建者的会话被占 124 s。在 t = 0、1、2、4、8、15、32、64 s 各探测一次,最后以FAILED workspace_unavailable结束(S21);期间后续 Turn 与其他变更都得到409 session_context_busy。 - T3: 获取路径上的
isReadable/isExecutable检查仍然只有真实栈能钉住。第 4 轮的candidate-acquire-modes.patch仍能干净地打上。
合并时的注意事项: 开放中的 #13260、#13354 也占用 V46,需要在最终的 merge ref 上重新检查唯一性。
暂时无法测试: #13265 的 child_run/monitor_run 在生产配置下仍未启用。启用后,后台 Shell 会比它所在的 Turn 活得更久,而 W2 的屏障只统计活跃 Turn、未结操作和已请求的 Action。届时这个握手、以及设计所依赖的「每个 Turn 都使用全新 Runtime Session」这一前提,都值得重新检查。这一点来自读代码,我没有实测。
证据:wenshao/qwen-code@3b0e111 → pr13247/r5/ · 第 4 轮:issuecomment-6007931423
|
@qwen-code /triage |
qqqys
left a comment
There was a problem hiding this comment.
Reviewed head: f583ec3774366f8424feef4f91ed85e594a7b044 (base main).
Approve. Every Critical filed across all four rounds was re-verified in source at this exact head and is fixed; the current Critical-only scan found nothing blocking.
Note on evidence: GitHub's reviews endpoint returned a server-side 422 for this PR throughout the review, so the historical findings were enumerated from the inline review-comment stream rather than from review bodies or ledger metadata. Eight distinct Criticals were identified and each was checked in current code, not from a resolution flag or a later round's silence.
Historical Criticals — status at this head
[R1-1] beginCwdChangeOperation ran its only authorization check after the replay lookup, and the service never called requireReadGrant — fixed. requireCwdChangeActor(session, actorId) now runs before the replay SELECT, and it delegates to requireWorkspaceCreator. So an actor whose managed_workspace_access.can_read was revoked after admitting key K fails the creator check on re-POST and never reaches the actor_digest/idempotency_key match that used to hand back the stored record. Settlement re-proves the grants rather than trusting admission: hasCwdChangeRegistryFacts joins the registry, create command and access rows and requires r.state = 'ACTIVE', matching generation and storage, and a.can_read = TRUE AND a.can_create = TRUE for the creation actor.
[R1-2] The migration claimed a Flyway version already taken on main, invisible from this branch — fixed. The file is V46__managed_cwd_operation.sql after two renumberings, and the Flyway migration version uniqueness lane passes at this head.
[R2-1] The deployment opt-in gate preceded the replay lookup, inverting the documented idempotency guarantee — fixed. The order is now actor check → replay → gate, and the comment states the reason: the replay follows the actor check so a lost 202 still resolves to the original operation even after the deployment has disabled execution.
[R2-2] The settlement probe verified the target against the Session's current cwd, terminally refusing a change away from a destroyed directory — fixed. verifyInstallable builds a candidate ContextBinding carrying WorkspaceRelativePath.normalize(targetCwdRelative) and runs authority.verifyMountForProbe on that candidate, so the guard no longer couples to the old directory's continued existence. The Javadoc names this as the escape the feature exists for.
[R2-3] The shared requireDirectory had no access check while the worker enforces R_OK | X_OK — fixed on both paths. The acquire-path requireDirectory now requires Files.isReadable(directory) && Files.isExecutable(directory) alongside containment, isDirectory(NOFOLLOW_LINKS) and realpath identity; the probe twin uses provider().checkAccess(directory, READ, EXECUTE) so a refusal reaches the classifier's catch arms instead of being swallowed by a false-returning predicate.
[R2-4] The bound later-Turn barrier reused the shared hasOpenOperation, whose counted population is wider than the contract promises, so a stuck PENDING rename could wedge Turn submission permanently — fixed. insertTurnCommand no longer calls hasOpenOperation; for a bound Session it calls the narrower hasOpenExecutionOperation, and the comment records the exact wedge this avoids — "a stuck display mutation (PENDING rename command) or an in-flight ACTION_RESPONSE would otherwise wedge every Turn admission without any recovery scan reading those tables" — while the operation-ledger barriers that intentionally cover those tables (requireNoOpenOperation, validateOperationStart) are left untouched.
[R3-1] The cwd busy barrier was the only admission barrier that did not consult hasDecidableAction, so a cwd change could be admitted and committed while a permission Action was still requested — fixed. Both cwd barrier sites now OR in hasDecidableAction(session), so a still-answerable requested Action blocks admission the same way an open operation does.
[R4-1] Failure classification was drawn by exception origin rather than momentary-vs-permanent; filed three times, the third adding an acquire-path regression — fixed, including the specific trap the third filing named. The resolver now carries the acquire/probe split the finding asked for: resolve() uses verifyMountIntact, whose single IOException arm yields the terminal verdict, and the shared requireDirectory catches IOException | IllegalArgumentException | SecurityException terminally, with a comment naming the regression it prevents ("an ENOTDIR-shaped cwd must not surface as hosted_harness_unavailable or park the binding RECOVERY_BLOCKED"). Only verifyInstallable uses the ...ForProbe twins, where a momentary fault may retry against the caller's budget. Critically, the fix does not key on NotDirectoryException | FileSystemLoopException — which the finding warned measures as a no-op on JDK 21, where ENOTDIR and ELOOP surface as a bare FileSystemException. Instead the probe's residual IOException arm calls hasStructuralAncestor(base, directory) and returns terminal when the ancestor chain holds a regular file or a dangling/looping symlink, leaving only genuinely opaque blips (ESTALE/EIO, ENAMETOOLONG) retryable.
Critical-only scan
Read in full: the resolver's acquire and probe classification paths, the store's cwd admission ordering, both busy-barrier predicates, hasOpenOperation and hasDecidableAction, and insertTurnCommand's barrier. No Critical found. The two properties this slice turns on both hold in code: a refused change cannot redirect tool execution, because the acquire path's directory and mount rules stay terminal and the probe only ever reports; and the barrier is narrow enough that a stuck display mutation cannot wedge the Session's only re-acquire path.
Not scanned — disclosed, not asserted clean. I did not audit the two controllers, ApiModels, the SessionLifecycleService/Coordinator bodies beyond the barrier call sites, the WorkspaceStorageGuard probe twins, the contract JSON and generated TypeScript client, the V46 SQL beyond its version, or the design docs. I report no Critical there because I found none where I looked, not because I proved absence.
CI
Fully settled and green at this head: 20 pass, 3 skipped, 0 failed, 0 pending, including the Flyway migration version uniqueness lane that R1-2 turned on. Nothing indicates a defect introduced by this PR.
Scope note
Approval is bound to commit f583ec37. This is a code-level review: no local Java or TypeScript suite execution, no real MySQL/MariaDB run, and no live multi-actor authorization probe against a running control plane.
|
Approve. Every Critical filed across all four rounds was re-verified in source at this exact head and is fixed; the current Critical-only scan found nothing blocking. |
yiliang114
left a comment
There was a problem hiding this comment.
Approve. Every Critical filed across all four rounds was re-verified in source at this exact head and is fixed; the current Critical-only scan found nothing blocking.
…#13565) Move the merged-PR history of the Managed Agent dual-path proposal (QwenLM#12380) out of the issue body into a bilingual ledger under docs/design/. The issue body had come within 10 KB of GitHub's 256 KiB limit, so the body will keep only the delivery snapshot and the open PRs, and merged rows move here rewritten to their merged final state. The ledger carries every merged row the issue tracked up to its 2026-10-03 reconcile (adding the missing merge commit to the five earliest rows and normalising the Chinese state cells), rewrites the eight rows the issue still listed as open although their PRs had merged (QwenLM#13141, QwenLM#13166, QwenLM#13174, QwenLM#13210, QwenLM#13214, QwenLM#13217, QwenLM#13218, QwenLM#13247), and adds rows for the 48 managed-agent PRs merged between that reconcile and main 0c13502 that had no row yet. PRs closed without merging (QwenLM#13087, QwenLM#13336) sit in their own table. Later merges land at the next reconcile. Co-authored-by: wenshao <[email protected]>








What this PR does
Implements the W2 slice of proposal #12380: a controlled working-directory change for a Workspace-bound Managed Session. The creator asks the control plane to move the Session's relative directory inside the same authorized Workspace, and the change runs as a durable, idempotent operation:
POST /v1/agents/sessions/{sessionId}/cwdtakes the Idempotency-Key header, a target directory and the expected context revision, answers 202 with acwd_changeoperation, and the WebShell twin/api/agent/web-shell/v1/sessions/cwd/changemirrors it with a body-carried key.Admission is an ordered pipeline: argument resolution (400s) and the trusted-actor check (401), the binding gate, then the creator check — an unreadable Session stays invisible at 404 while a readable non-creator hears
403 session_operation_forbidden, the exact pairing the merged bound lifecycle (#13135, #13194) ships — then the replay lookup (same key and normalized payload returns the original operation even after it completes, so a lost 202 resolves even when execution has since been disabled), the deployment opt-in gate, active-status and registry-fact checks, a revision CAS (409 context_revision_conflict), and the busy barrier (409 session_context_busywhile a Turn is active or any operation is open). A background coordinator claims the admitted operation, verifies the target against the administrator mount table (canonical path, filesystem identity, the storage guard when verified recovery is enabled, and the exact directory rule a tool-turn acquisition enforces), and commits in one transaction:cwd_relativeandcontext_revision = expected + 1update, operation completion, and asession.context.changedevent. A probe refusal or a fact that moved between claim and commit ends the operation asfailedwith a typedfailure_codeand never retries; the Session provably keeps its prior directory and revision, so a refused change can never redirect tool execution. The next tool turn after a commit acquires a fresh Runtime Session and installs exactly the committed binding, so no worker or Harness protocol change was needed. Since #13112 landed meanwhile, this PR also closes the documented handshake in the other direction: bound later-turn submission refuses409 session_context_busywhile a context-changing operation (a cwd change or a lifecycle operation) is open — deliberately without blocking on stuck display mutations or in-flight permission Actions, which would otherwise wedge the Session's only re-acquire path.The reviewed contract moves from v1.29.0 to v1.32.0 (v1.29.0 shipped with #13142; #13112 has since landed and recorded its later Turns as v1.28 in the upstream header, #13210 landed v1.30.0 with its gateway-free qwenSignature scheme, and #13265 landed v1.31.0 with the Stage H3 background task journal): the two routes and their schemas flip from
plannedtoimplemented, the operation query serves the cwd shape on both surfaces,failedoperations are required to carry a non-nullfailure_code, and the operation status enum ships only the four values the server can emit. The design is recorded in two linked documents that must stay synchronized: English and 简体中文.Why it's needed
Today a bound Session pins its directory at creation and the only way out of a bad choice is abandoning the conversation. W2 is the smallest claimable unowned slice on the tracker ("Controlled directory-change admission and settlement; public/WebShell routes remain
planned"), and everything it needs was already merged: the binding withcontextRevision, the durable operation ledger, the Workspace turn machinery, the documentation for actor-scoped idempotent operations. The slice deliberately stays narrow: cross-Workspace moves and storage migration are new Sessions by design; the Session read model keepsWorkspaceContext.state: ready(the operation and the event are the observation channel); settlement verifies deterministically on the same host and mount the worker uses instead of claiming storage for a receipt probe — a crash can strand nothing. Where the reference contract demanded cross-process receipts, the merged architecture already hands the first post-change turn a fresh install with the worker's own receipt before any tool runs, in the correct directory and typed (managed_context_unavailable) if it has meanwhile become unusable.Reviewer Test Plan
How to verify
cd packages/sdk-java/managed-agent-server && mvn clean verify checkstyle:check— the full suite (1009 tests incl. the new ones) passes on H2 with the SpotBugs gate green. Focused classes:ManagedCwdChangeOperationTest(32 tests: admission matrix incl. cross-tenant invisibility, actor-refusal precedence and probing invisibility of revoked grants, replay over PENDING/COMPLETED/FAILED and across a deployment-flag flip, revision CAS, busy barriers incl. the two-way race, the later-turn handshake in both directions, barrier release in both terminals, same-directory change, per-conjunct claim fences, stale-owner guards, coordinator settle/refuse/retry/reclaim through the recovery scan itself),WorkspaceRuntimeInstallProbeTest(7: symlink/alias/missing/./fileKey swap, candidate-binding guard, unreadable target),WorkspaceRuntimeResolutionPivotTest(the post-commit pivot),ManagedOperationSchemaUpgradeTest(additive columns read tolerantly under a pinned V31 schema, with a real row exercised),ManagedAgentApiContractTest(no contract drift; both routes exercised for 202/400/401/404/409).mvn test -Dtest='HostedPublicWorkspaceIT' -Dnode.executable=$(which node)— 3 tests boot the real server, Broker, Hosted Harness (dist/cli.js) and a fixture model: create a bound Session, change its directory through both surfaces with normalized idempotent replays, refuse missing directories with a terminalfailedoperation andfailure_code, recover with another change, count exactly onesession.context.changedper completed change (three, including the WebShell one and the recovery), then submit a bound later Turn (merged feat(managed-agent): let a Workspace-bound Session's creator submit, cancel and rename #13112) and prove its file writes land only in the committed directory.replayed=trueeven spelled differently (child2//vschild2/./); a stale expected revision answers 409context_revision_conflict; a stranger always gets 404 while a readable non-creator gets 403session_operation_forbidden; afailedoperation reportsfailure_codeand the Session still serves reads and later changes.ManagedAgentApiContractTestasserts no mapped-but-planned drift and pins the two routes plus four DTO schemas against the regenerated WebShell client types.Evidence (Before & After)
N/A (non-UI; the WebShell BFF route is consumed by a later UI slice by design). Test evidence and the E2E result matrix are appended in the plan file
.qwen/e2e-tests/managed-workspace-w2-cwd-change.md(also posted as a PR comment).Tested on
Environment (optional)
JDK 21, H2
MODE=MySQL, Node 22, dist/cli.js as Hosted Harness; MySQL CI lanes cover the hosted MySQL profile (not run locally — CI covers it).Risk & Scope
access(R_OK|X_OK)rule, so a change toward a destroyed or unreadable directory is refused typed at settlement time instead of wedging the first post-change turn (the design documents why the probe cannot strand storage on a crash, unlike the claim-based turn path).WorkspaceContext.statederivation in reads; model-context notes; WebShell UI; cross-Workspace moves and storage migration; verified-recovery MySQL/MariaDB fault matrices and second-host recovery.Follow-ups recorded on the tracker: two recorded review deferrals — the WebShell
@Size(128)over-length key classifying asinvalid_requestwhile the public surface usesinvalid_idempotency_key(disclosed in the WebShell route description), and a distinct wire code for the retryable probe classification (the terminal/transient split is pinned in both directions and logged with its cause, but exposing it as e.g.workspace_mount_unverifiablechanges the published 409 vocabulary and the hosted tool-turn classifier's code set atomically); the worker-receipt probe for cross-host runtimes; andrecovery_blockedre-admission once a producer exists.Linked Issues
Refs #12380 (implements the W2 row of the delivery snapshot. Does not close the proposal.)
中文说明
本次 PR 做了什么
实现提案 #12380 的 W2 切片:为绑定 Workspace 的 Managed Session 提供受控目录变更。创建者请求控制面把会话的相对目录移到同一授权 Workspace 内的另一个位置,变更以持久化、幂等的 operation 运行:公开路由
POST /v1/agents/sessions/{sessionId}/cwd接收 Idempotency-Key 头、目标目录与期望上下文 revision,答 202 返回cwd_changeoperation;WebShell 孪生路由/api/agent/web-shell/v1/sessions/cwd/change以请求体携带键镜像。准入按固定管线执行:参数解析(各 400)与可信 actor 检查(401)、绑定门槛,随后创建者校验——无读授权的会话保持不可见答 404,有读授权的非创建者答
403 session_operation_forbidden,与已合入的绑定生命周期(#13135、#13194)发布的形态一致——然后重放查找(同键且归一化后的载荷相同则返回原 operation,即使它已完成;执行随后被禁用时,丢失的 202 也能经原 operation 解析)、部署开关门槛、活动状态与 Registry 事实、revision CAS(冲突时 409context_revision_conflict)以及繁忙屏障(存在活动 Turn 或未关闭 operation 时 409session_context_busy)。后台协调器认领 operation,对照管理员挂载表核验目标(规范路径、文件系统身份、已启用验证恢复时还包括存储守卫,以及工具轮次获取时执行的同一目录规则),然后在单事务中提交:更新cwd_relative与context_revision = expected + 1、标记 operation 完成、追加session.context.changed事件。探针拒绝或认领与提交之间事实发生变化,operation 以类型化的failure_code终态失败且绝不重试;会话可证明地保持原目录与 revision,因此一次被拒绝的变更绝不会让工具跑进错误目录。提交后的下一个工具轮次在新的 Runtime Session 上恰好安装已提交的绑定,因此无需修改 worker 或 Harness 协议。由于 #13112 在此期间合入,本 PR 同时反向补上文档记录的衔接:存在未关闭的上下文变更型 operation(cwd 变更或生命周期 operation)时,绑定会话的后续 Turn 准入拒绝409 session_context_busy——但不会被卡死的展示变更或在途的权限 Action 楔死,因为 Turn 准入是会话唯一的重获通道。已评审契约从 v1.29.0 升至 v1.32.0(v1.29.0 已随 #13142 发布;#13112 随后合入并在上游头信息中把后续 Turn 记为 v1.28,随后 #13210 以网关免鉴权 qwenSignature 方案发布 v1.30.0、#13265 以 H3 后台任务 journal 发布 v1.31.0):两条路由与其 schema 从
planned翻转为implemented,operation 查询在两侧都返回 cwd 形态,failed状态的 operation 必须携带非空failure_code,operation 状态枚举只发布服务器实际能发出的四个值。设计记录在两份必须保持同步的文档中:English 与 简体中文。为什么需要
今天绑定会话的目录在创建即固定,一次选错只能放弃整段对话。W2 是跟踪台账上无人认领的最小切片("受控目录变更准入与结算;公开/WebShell 路由仍为
planned"),它所需的底座早已合入:带contextRevision的绑定、持久化 operation 账本、Workspace 轮次机制、actor 域幂等操作的录入。切片刻意收窄:跨 Workspace 移动与存储迁移按设计新建会话;读模型保持WorkspaceContext.state: ready(operation 与事件即观察通道);结算在与 worker 相同的主机与挂载上做确定性核验,而不是为了一个回执去占用存储——崩溃不会遗留任何需要人工恢复的租约。参考契约要求跨进程回执的地方,合入架构已通过"提交后首个 Turn 经生产路径安装并获得 worker 自己的回执"提供同样的正确目录保证,且期间若目录失效会以类型化错误受阻。评审测试计划
如何验证
cd packages/sdk-java/managed-agent-server && mvn clean verify checkstyle:check— 全量套件(1009 个测试)在 H2 全绿,SpotBugs 门禁通过。重点类:ManagedCwdChangeOperationTest(32 个:准入矩阵含跨租户不可见、actor 拒绝次序与撤权重放不可见、三态及部署开关翻转下的重放、revision CAS、双竞态、后续 Turn 双向衔接、两种终态下的屏障释放、同目录变更、逐项认领栅栏、陈旧属主守卫、经恢复扫描自身的协调器结算/拒绝/重试/回收),WorkspaceRuntimeInstallProbeTest(7 个:软链/别名/缺失/./fileKey 交换、候选绑定的守卫、不可读目标),WorkspaceRuntimeResolutionPivotTest(提交后枢轴),ManagedOperationSchemaUpgradeTest(固定 schema 下有真实行时增量列读取宽容),ManagedAgentApiContractTest(无契约漂移;202/400/401/404/409 行为钉)。mvn test -Dtest='HostedPublicWorkspaceIT' -Dnode.executable=$(which node)— 3 个测试拉起真实 server、Broker、Hosted Harness(dist/cli.js)与 fixture 模型:创建绑定会话、经两侧 API 变更目录并做归一化幂等重放、对缺失目录以终态failed+failure_code拒绝、再以新变更恢复,断言每次完成恰有一条session.context.changed(共三条,含 WebShell 与恢复),随后提交一个绑定后续 Turn(已合入的 feat(managed-agent): let a Workspace-bound Session's creator submit, cancel and rename #13112)并证明其文件写入只落在已提交的目录。child2//vschild2/./)重放也返回原 operation id 且replayed=true;过期期望 revision 答 409context_revision_conflict;陌生人永远 404,有读授权的非创建者得 403session_operation_forbidden;failed的 operation 带failure_code且会话可继续读取与再次变更。ManagedAgentApiContractTest断言无 mapped-but-planned 漂移,并把两条路由与四个 DTO schema 钉在再生成的 WebShell 客户端类型上。证据(前后对比)
N/A(非 UI;WebShell BFF 供后续 UI 切片使用)。E2E 结果矩阵在
.qwen/e2e-tests/managed-workspace-w2-cwd-change.md(随 PR 评论附上)。验证环境
运行环境(可选)
JDK 21、H2
MODE=MySQL、Node 22、用 dist/cli.js 作为 Hosted Harness;MySQL CI 车道覆盖 hosted MySQL 形态(本机未跑,由 CI 覆盖)。风险与范围
access(R_OK|X_OK)规则,因此目标目录被摧毁或不可读的变更在结算期就以类型化拒绝,而不是楔死提交后的首个 Turn(设计文档解释了探针为何不会像 turn 路径那样在崩溃时留置存储租约)。WorkspaceContext.state派生;模型上下文说明;WebShell UI;跨 Workspace 移动与存储迁移;MySQL/MariaDB 故障矩阵与第二主机恢复。跟踪台账记录的后续项:两项评审延期——WebShell 的
@Size(128)超长键判invalid_request而公开面判invalid_idempotency_key(已在 WebShell 路由描述中披露),以及为可重试探针分类单独设立 wire 错误码(终态/瞬时分类已在两个方向钉住并带因记日志,但将其暴露为如workspace_mount_unverifiable会同时改动已发布的 409 词汇表与 hosted tool-turn 分类器的代码集合);面向跨主机运行时的 worker 回执探针;以及出现产生者后重新接纳recovery_blocked。关联 Issue
Refs #12380(实现交付快照中的 W2 行;不关闭提案。)