Skip to content

feat(managed-agent): admit workspace-bound sessions without execution - #12709

Merged
doudouOUC merged 9 commits into
mainfrom
codex/managed-workspace-w0b
Sep 26, 2026
Merged

doudouOUC merged 9 commits into
mainfrom
codex/managed-workspace-w0b

Conversation

@doudouOUC

Copy link
Copy Markdown
Collaborator

What this PR does

Adds the W0b admission slice for Managed Agent Sessions. A caller can select a registered Workspace and relative directory when creating an empty Session through the public or WebShell API. The Session persists the seven-field binding and frozen configuration references, and subsequent reads honor the actor's current Workspace read grant. Idempotent retries return the original binding even after Registry changes. Bound Turn execution and lifecycle mutations remain closed at the API, persistence, recovery, and embedded Runtime boundaries.

Why it's needed

The exploratory stack in #12358 demonstrates the end-to-end direction, but admitting a Workspace must be independently reviewable before any Worker mount or execution path is enabled. This PR builds on the merged Spring control plane in #12692 and keeps legacy unbound Sessions unchanged.

Reviewer Test Plan

How to verify

  • With a trusted tenant-scoped actor and a registered Workspace grant, create an empty Session using an explicit Workspace and a relative directory containing redundant . segments. Expect an accepted response with the selected Workspace and normalized directory, but no Turn or Runtime dispatch.
  • Repeat the same idempotency key after changing the Workspace default, generation, or state. Expect the original Session and binding; changing the request content should produce an idempotency conflict.
  • Revoke the read grant, then get or list the Session, read its events/items/transcript, or retry creation. Expect no bound Session data to be returned. A missing or mismatched trusted actor must not gain access through headers or request JSON.
  • Attempt non-empty bound creation, a bound Turn or lifecycle operation, and a recovered bound Turn. Expect a closed path without a legacy Hosted Harness or global Runtime call. Confirm unbound Sessions still follow the existing path.

Evidence (Before & After)

Before: the Spring API could only create legacy unbound Sessions. After: it admits metadata-only Workspace-bound Sessions while explicitly refusing bound execution. This is an API/storage change; no TUI screenshots apply.

Tested on

OS Status
🍏 macOS ✅ 68 Java unit/integration tests, 6 MySQL integration tests, Checkstyle, repository build and typecheck
🪟 Windows ⚠️ not tested locally
🐧 Linux ⚠️ not tested locally

Environment (optional)

Java 21, Maven 3.9.16, H2, ephemeral MySQL 8.4, Node.js 22 workspace build. The MySQL verification used a fresh schema; a separate focused SSE revocation regression also passed after the full run.

Risk & Scope

  • Main risk or tradeoff: Registry and grants are administrator-populated, and the standalone server has no production trusted-actor adapter. The host must supply a verified principal before real HTTP admission is usable.
  • Not validated / out of scope: bound Worker mounts, physical isolation, Hosted Harness execution, and production authenticated HTTP E2E. These remain W0c gates; this PR must not be treated as execution-ready.
  • Breaking changes / migration notes: three additive managed-agent SQL migrations store the binding, Registry/grants, and creation receipt. Existing unbound Session responses and placement remain unchanged.
  • Design: English · 简体中文.

Linked Issues

Related exploratory PR: #12358. Builds on merged #12692.

中文说明

本 PR 做了什么

实现 Managed Agent Session 的 W0b 准入切片。调用方可通过公开或 WebShell API,在创建空 Session 时选择已注册 Workspace 和相对目录。Session 持久化七字段绑定与冻结的配置引用;后续读取检查 actor 当前的 Workspace 读权限。Registry 改变后,幂等重试仍返回原绑定。绑定 Turn 的执行和生命周期修改在 API、持久化、恢复调度与嵌入式 Runtime 边界继续关闭。

为什么需要

#12358 的探索性完整方案展示了端到端方向,但 Workspace 准入应在启用任何 Worker 挂载或执行路径前独立评审。本 PR 基于已合入的 Spring 控制面 #12692,保持未绑定旧 Session 的行为不变。

评审测试计划

如何验证

  • 使用可信的租户级 actor 和已注册 Workspace 授权,显式选择 Workspace,并以带冗余 . 段的相对目录创建空 Session。应返回所选 Workspace 和规范化目录,但不创建 Turn,也不派发 Runtime。
  • 改变默认 Workspace、generation 或状态后,使用相同幂等键重试。应返回原 Session 和绑定;改变请求内容应得到幂等冲突。
  • 撤销读权限后,获取或列出 Session、读取事件/item/transcript,或重试创建。不能返回绑定 Session 数据。缺失或租户不匹配的可信 actor 不能靠 header 或请求 JSON 获得权限。
  • 尝试非空绑定创建、绑定 Turn 或生命周期操作,以及恢复已持久化的绑定 Turn。旧 Hosted Harness 和全局 Runtime 路径均不得被调用;未绑定 Session 应继续走原有路径。

前后证据

之前:Spring API 只能创建未绑定的旧 Session。之后:可准入仅有元数据的 Workspace 绑定 Session,同时明确拒绝绑定执行。这是 API/存储改动,不适用 TUI 截图。

本地验证平台

系统 状态
🍏 macOS ✅ 68 个 Java 单元/集成测试、6 个 MySQL 集成测试、Checkstyle、仓库构建及类型检查
🪟 Windows ⚠️ 本地未测试
🐧 Linux ⚠️ 本地未测试

环境

Java 21、Maven 3.9.16、H2、临时 MySQL 8.4 和 Node.js 22 工作区构建。MySQL 验证使用全新 schema;完整测试后另行通过了 SSE 权限撤销的定向回归测试。

风险与范围

  • 主要风险或权衡:Registry 和授权由管理员填充;独立服务没有生产级可信 actor 适配器。宿主必须提供经验证的 principal,真实 HTTP 准入才能使用。
  • 未验证/不在范围内:绑定 Worker 挂载、物理隔离、Hosted Harness 执行,以及生产认证 HTTP E2E。这些仍是 W0c 门禁;不能将本 PR 当作执行就绪。
  • 兼容性与迁移:三个增量 managed-agent SQL 迁移保存绑定、Registry/授权与创建回执。现有未绑定 Session 的响应和放置行为不变。
  • 设计文档:English · 简体中文。

关联事项

相关探索 PR:#12358。基于已合入的 #12692。

@doudouOUC

Copy link
Copy Markdown
Collaborator Author

E2E test report (W0b)

Status: production authenticated HTTP E2E not run. The standalone Spring server does not include a production adapter that supplies a verified tenant-scoped actor, and W0b intentionally refuses bound Turn execution. The global qwen CLI was dry-run at version 0.24.4, but it does not expose this standalone managed-agent API, so it cannot serve as a bound-Session baseline.

Narrower verification completed on macOS: 68 Java unit/H2 integration tests, 6 fresh-schema MySQL 8.4 integration tests including cross-JVM retry and exact-case isolation, a focused SSE revocation test, Checkstyle, repository build, and typecheck. MockMvc injected a trusted principal to exercise public and WebShell admission, replay, revocation, and execution rejection; coordinator and Broker tests verified fail-closed behavior. This evidence does not prove a production authentication adapter, Worker mount isolation, or bound execution. Those remain separate W0c gates.

The local E2E plan is recorded under .qwen/e2e-tests/2026-09-26-managed-workspace-admission.md (ignored working artifact).

@wenshao

wenshao commented Sep 25, 2026

Copy link
Copy Markdown
Collaborator

Maintainer verification: real MySQL + real server JAR, head cd76773f

Verdict: the W0b behavior is correct end to end (62/62 live checks). I recommend fixing F1 before merge. F1 is the measured answer to the CAST question that triage left open (stage 2, stage 3). The rewrite does take the legacy unbound paths off their indexes, and the FOR UPDATE variant locks every tenant's command rows. A +29/−17 candidate patch keeps the exact-byte check and restores main's query plans. I verified it below. F2 (three guards with no test) and F3 (Idempotency-Key namespaces) are smaller and do not block merge.

What I ran

  • Environment: macOS arm64, JDK 21.0.12, Maven 3.9.16, MySQL 8.4.7 on a fresh datadir. Control arm: origin/main 9e51e6ac, whose packages/sdk-java is identical to the PR base be1a0a1f.
  • Suites: mvn verify on managed-agent-server: 69/69 passed, Checkstyle clean. -Pmysql-integration: 6/6 on a fresh default-collation database, and 6/6 on a utf8mb4_unicode_ci database (the collation behind the feat(managed-agent): Spring control plane and dual-path WebShell #12692 F1). All four new tables and the new session columns come out as utf8mb4_bin.
  • Live rig:
    • The real Spring Boot JAR, started through PropertiesLauncher.
    • A 60-line adapter JAR that sets an AuthenticatedTenantActor principal from X-Rig-Actor. It stands in for the host's verified ingress, which the standalone server does not have. By construction it trusts that header, so it does not replace production-auth E2E.
    • A counting fake Hosted Harness that logs every request and answers 503.
    • Registry rows, grants and the tenant default seeded by an administrator over SQL.

Reviewer Test Plan, live (Fig. 1)

All 62 checks passed:

  • Admission normalized ./services/./api/ to services/api and stored the seven-field binding plus the frozen configRef/policyRef. No Turn row was written and the Harness got no request.
  • A retry after generation=2, DRAINING and a default change returned the original Session and binding. A changed cwd returned 409 idempotency_conflict.
  • The actor is taken only from the principal. Header and JSON-body claims got 401. A tenant mismatch got 403. Alice and ACME case variants got 404.
  • All 10 malformed selections were rejected before any row was written.
  • Every bound Turn and lifecycle operation got 409 for a reader and 404 for a non-reader. No command or Turn row was written, and the Harness got no request.
  • I inserted a bound ACCEPTED Turn directly in MySQL to simulate a persisted one. The dispatcher marked it FAILED workspace_unavailable with the Harness untouched.
  • Revocation returned 404 on every read route and hid the Session from lists. An open SSE stream closed about 2 s after the grant was revoked.
  • The legacy unbound path still reaches the Harness. This is the positive control that the counter works.

live walk-through

F1 (fix before merge): the CAST(CONCAT(col,'!') AS BINARY(513)) predicates are not index-usable on MySQL (Fig. 2)

Setup: 100,000 legacy Sessions across 1,000 tenants, 100,000 command rows and 20,000 grants, with the same data for each arm.

  • Plans:

    • listSessions goes from a ref on managed_agent_session_list_idx (100 rows) to a full scan of 99,375 rows plus a filesort. Its new grant EXISTS scans all 20,000 grants for each bound row.
    • findCommand … FOR UPDATE goes from a const PK lookup to a full scan of 98,670 rows. canRead is also a full scan.
    • findSession/requireSessionForUpdate are unaffected, because UNIQUE(session_id) still serves them. Triage called that correctly.
  • Locks: under REPEATABLE READ, the PR's findCommand … FOR UPDATE for tenant t0500 takes 101,650 record locks (main takes 1). Meanwhile tenant t0001's own lookup and its INSERT into managed_agent_command both fail with ERROR 1205 Lock wait timeout. That statement runs inside insertTurnCommand, beginSessionMutation and completeSessionMutation. As a result, legacy Turn submission and every lifecycle mutation serialize across all tenants.

  • Real HTTP latency (median):

    Path main PR Slowdown
    List 6.9 ms 88.9 ms 13×
    Create 5.8 ms 46.9 ms 8×
    16 tenants × 10 Turn submissions in parallel 39 ms (0.7 s wall) 687 ms (7.0 s wall) 17×
    Bound list, 300 bound Sessions — 687 ms —

    All of these scale linearly with table size, and the command table only grows.

  • What the CAST buys: it is not a no-op. utf8mb4_bin is PAD SPACE on MySQL 8.4, so 'ws-a' = 'ws-a ' is TRUE with a plain =. So it adds exactness for Workspace IDs. For tenant IDs and idempotency keys it has no effect, because their validators already reject spaces.

  • Candidate fix: put a sargable col = ? in front of the CAST predicate in 7 queries, and keep the CAST. Any row the CAST accepts also satisfies the new =, so results are unchanged.

    • Verified: 69/69 unit/H2, 6/6 MySQL IT, Checkstyle clean, 62/62 live walk-through.
    • Main's plans are restored: ref list_idx for the list, a const PK lookup for findCommand, and eq_ref/const PK for the grant lookups.
    • The lock test is back to 1 lock with no cross-tenant wait. Latency: list 7.6 ms, create 5.1 ms, parallel submit 24 ms median / 0.4 s wall, bound list 9.5 ms.

mysql index regression

Candidate patch for F1 (+29/−17, ManagedAgentStore + ManagedWorkspaceRegistry)
--- a/packages/sdk-java/managed-agent-server/src/main/java/com/alibaba/qwen/code/managedagent/store/ManagedAgentStore.java
+++ b/packages/sdk-java/managed-agent-server/src/main/java/com/alibaba/qwen/code/managedagent/store/ManagedAgentStore.java
@@ -217,9 +217,10 @@ public class ManagedAgentStore implements AgentStateStore {
             String actorId, String idempotencyKey) {
         return jdbc.query("SELECT request_digest, session_id, turn_id"
                         + " FROM managed_workspace_create_command"
-                        + " WHERE CAST(CONCAT(tenant_id, '!') AS BINARY(513))"
+                        + " WHERE tenant_id = ? AND actor_id = ?"
+                        + " AND idempotency_key = ?"
+                        + " AND CAST(CONCAT(tenant_id, '!') AS BINARY(513))"
                         + " = CAST(CONCAT(?, '!') AS BINARY(513))"
-                        + " AND actor_id = ?"
                         + " AND CAST(CONCAT(idempotency_key, '!') AS BINARY(513))"
                         + " = CAST(CONCAT(?, '!') AS BINARY(513))",
                 (result, row) -> new WorkspaceCommand(
@@ -227,7 +228,7 @@ public class ManagedAgentStore implements AgentStateStore {
                         result.getString("session_id"),
                         result.getString("turn_id")),
                 tenantId, ManagedWorkspaceRegistry.actorKey(tenantId, actorId),
-                idempotencyKey);
+                idempotencyKey, tenantId, idempotencyKey);
     }
 
     private record WorkspaceCommand(String requestDigest, String sessionId,
@@ -507,10 +508,11 @@ public class ManagedAgentStore implements AgentStateStore {
                         + " command_status, session_status_before,"
                         + " created_at, updated_at"
                         + " FROM managed_agent_command WHERE"
+                        + " tenant_id = ? AND operation = ?"
+                        + " AND idempotency_key = ? AND"
                         + " CAST(CONCAT(tenant_id, '!') AS BINARY(513))"
                         + " = CAST(CONCAT(?, '!') AS BINARY(513))"
-                        + " AND operation = ? AND"
-                        + " CAST(CONCAT(idempotency_key, '!') AS BINARY(513))"
+                        + " AND CAST(CONCAT(idempotency_key, '!') AS BINARY(513))"
                         + " = CAST(CONCAT(?, '!') AS BINARY(513))"
                         + (forUpdate ? " FOR UPDATE" : ""),
                 (result, row) -> new CommandRecord(
@@ -524,7 +526,8 @@ public class ManagedAgentStore implements AgentStateStore {
                         result.getString("session_status_before"),
                         result.getLong("created_at"),
                         result.getLong("updated_at")),
-                tenantId, operation, idempotencyKey);
+                tenantId, operation, idempotencyKey, tenantId,
+                idempotencyKey);
         return rows.stream().findFirst();
     }
 
@@ -551,6 +554,7 @@ public class ManagedAgentStore implements AgentStateStore {
             Long beforeUpdatedAt, String beforeSessionId, int limit) {
         List<Object> arguments = new ArrayList<>();
         arguments.add(tenantId);
+        arguments.add(tenantId);
         arguments.add(actorId == null ? null
                 : ManagedWorkspaceRegistry.actorKey(tenantId, actorId));
         String cursorClause = "";
@@ -563,13 +567,17 @@ public class ManagedAgentStore implements AgentStateStore {
         }
         arguments.add(limit + 1);
         List<SessionRecord> rows = jdbc.query(
-                "SELECT * FROM managed_agent_session WHERE"
-                        + " CAST(CONCAT(tenant_id, '!') AS BINARY(513))"
+                "SELECT * FROM managed_agent_session WHERE tenant_id = ?"
+                        + " AND CAST(CONCAT(tenant_id, '!') AS BINARY(513))"
                         + " = CAST(CONCAT(?, '!') AS BINARY(513))"
                         + " AND status <> 'DELETED'"
                         + " AND (workspace_id IS NULL OR EXISTS (SELECT 1"
                         + " FROM managed_workspace_access wa"
-                        + " WHERE CAST(CONCAT(wa.tenant_id, '!') AS BINARY(513))"
+                        + " WHERE wa.tenant_id = managed_agent_session.tenant_id"
+                        + " AND wa.workspace_id"
+                        + " = managed_agent_session.workspace_id"
+                        + " AND wa.actor_id = ?"
+                        + " AND CAST(CONCAT(wa.tenant_id, '!') AS BINARY(513))"
                         + " = CAST(CONCAT("
                         + "managed_agent_session.tenant_id, '!')"
                         + " AS BINARY(513))"
@@ -577,7 +585,6 @@ public class ManagedAgentStore implements AgentStateStore {
                         + " = CAST(CONCAT("
                         + "managed_agent_session.workspace_id, '!')"
                         + " AS BINARY(513))"
-                        + " AND wa.actor_id = ?"
                         + " AND wa.can_read = TRUE))"
                         + cursorClause
                         + " ORDER BY updated_at DESC, session_id DESC LIMIT ?",
--- a/packages/sdk-java/managed-agent-server/src/main/java/com/alibaba/qwen/code/managedagent/store/ManagedWorkspaceRegistry.java
+++ b/packages/sdk-java/managed-agent-server/src/main/java/com/alibaba/qwen/code/managedagent/store/ManagedWorkspaceRegistry.java
@@ -38,12 +38,14 @@ public class ManagedWorkspaceRegistry {
             return false;
         }
         return !jdbc.queryForList("SELECT 1 FROM managed_workspace_access"
-                + " WHERE CAST(CONCAT(tenant_id, '!') AS BINARY(513))"
+                + " WHERE tenant_id = ? AND workspace_id = ?"
+                + " AND CAST(CONCAT(tenant_id, '!') AS BINARY(513))"
                 + " = CAST(CONCAT(?, '!') AS BINARY(513))"
                 + " AND CAST(CONCAT(workspace_id, '!') AS BINARY(513))"
                 + " = CAST(CONCAT(?, '!') AS BINARY(513))"
                 + " AND actor_id = ? AND can_read = TRUE",
-                Integer.class, tenantId, workspaceId, key).isEmpty();
+                Integer.class, tenantId, workspaceId, tenantId, workspaceId,
+                key).isEmpty();
     }
 
     public ResolvedBinding resolveForCreation(String tenantId,
@@ -67,9 +69,10 @@ public class ManagedWorkspaceRegistry {
         if (selection == null) {
             List<String> defaults = jdbc.queryForList(
                     "SELECT workspace_id FROM managed_workspace_default"
-                            + " WHERE CAST(CONCAT(tenant_id, '!') AS BINARY(513))"
+                            + " WHERE tenant_id = ?"
+                            + " AND CAST(CONCAT(tenant_id, '!') AS BINARY(513))"
                             + " = CAST(CONCAT(?, '!') AS BINARY(513)) FOR UPDATE",
-                    String.class, tenantId);
+                    String.class, tenantId, tenantId);
             if (defaults.isEmpty()) {
                 throw workspaceRequired();
             }
@@ -81,12 +84,13 @@ public class ManagedWorkspaceRegistry {
                 "SELECT workspace_generation, storage_id, display_name,"
                         + " config_ref, policy_ref, state FROM"
                         + " managed_workspace_registry WHERE"
+                        + " tenant_id = ? AND workspace_id = ? AND"
                         + " CAST(CONCAT(tenant_id, '!') AS BINARY(513))"
                         + " = CAST(CONCAT(?, '!') AS BINARY(513))"
                         + " AND CAST(CONCAT(workspace_id, '!') AS BINARY(513))"
                         + " = CAST(CONCAT(?, '!') AS BINARY(513)) FOR UPDATE",
                 (result, row) -> workspaceRow(tenantId, workspaceId,
-                        result), tenantId, workspaceId);
+                        result), tenantId, workspaceId, tenantId, workspaceId);
         if (records.isEmpty()) {
             if (selection == null) {
                 throw workspaceRequired();
@@ -96,14 +100,15 @@ public class ManagedWorkspaceRegistry {
         }
         List<AccessRow> access = jdbc.query(
                 "SELECT can_read, can_create FROM managed_workspace_access"
-                        + " WHERE CAST(CONCAT(tenant_id, '!') AS BINARY(513))"
+                        + " WHERE tenant_id = ? AND workspace_id = ?"
+                        + " AND CAST(CONCAT(tenant_id, '!') AS BINARY(513))"
                         + " = CAST(CONCAT(?, '!') AS BINARY(513))"
                         + " AND CAST(CONCAT(workspace_id, '!') AS BINARY(513))"
                         + " = CAST(CONCAT(?, '!') AS BINARY(513))"
                         + " AND actor_id = ? FOR UPDATE",
                 (result, row) -> new AccessRow(result.getBoolean("can_read"),
                         result.getBoolean("can_create")), tenantId,
-                workspaceId, key);
+                workspaceId, tenantId, workspaceId, key);
         if (access.isEmpty() || !access.getFirst().canRead()) {
             if (selection == null) {
                 throw workspaceRequired();

F2 (recommended, test-only): three load-bearing guards have no test (Fig. 3)

I removed one guard at a time on a copy of cd76773f and ran the PR's 66 unit/H2 tests; the MySQL IT was not part of these runs. 20 of 28 mutants were killed.

  • Defense in depth (4 survivors): M02, M09, M11 and M18. An upstream guard still catches each one. For example, the store's non-empty check keeps the 409 when the service check is removed.
  • Untested integrity check (1): M23, the re-derivation of the frozen descriptor. It works live: a tampered row fails closed.
  • Load-bearing (3): I built a server JAR for each and ran it against a fresh MySQL. Each mutant changes only its own row; the PR code is correct on all three.
    • M12: after the read grant is revoked, a WebShell create retry gets 202 {sessionId: <original>, replayed: true} instead of 404. The public route stays 404 only because its controller re-reads the Session.
    • M14: the same key sent with ws-b/docs instead of ws-a/services/api gets 202 replay=true and silently keeps the original binding, instead of 409.
    • M24: a read-only grant creates a bound Session (202 instead of 403).

Three candidate tests (+74 lines in ManagedWorkspaceAdmissionTest) pass on the PR code (66 → 69) and kill M12, M14 and M24.

Candidate tests for F2 (+74 lines)
    @Test
    void webShellRetryAfterRevocationReturnsNoSession() throws Exception {
        String tenant = "tenant-" + UUID.randomUUID();
        register(tenant, "ws-a", "storage-a");
        grant(tenant, "ws-a", "actor-a", true);
        String body = """
                {"idempotencyKey":"web-retry","agentId":"qwen-code",
                 "input":[],"workspace":{"workspaceId":"ws-a"}}
                """;
        mvc.perform(post("/api/agent/web-shell/v1/sessions/create")
                        .header(TenantContextFilter.HEADER, tenant)
                        .principal(actor(tenant, "actor-a"))
                        .contentType(MediaType.APPLICATION_JSON)
                        .content(body))
                .andExpect(status().isAccepted());
        jdbc.update("UPDATE managed_workspace_access SET can_read = FALSE"
                        + " WHERE tenant_id = ? AND workspace_id = ?"
                        + " AND actor_id = ?", tenant, "ws-a",
                "actor-a".getBytes(java.nio.charset.StandardCharsets.UTF_8));
        mvc.perform(post("/api/agent/web-shell/v1/sessions/create")
                        .header(TenantContextFilter.HEADER, tenant)
                        .principal(actor(tenant, "actor-a"))
                        .contentType(MediaType.APPLICATION_JSON)
                        .content(body))
                .andExpect(status().isNotFound())
                .andExpect(jsonPath("$.sessionId").doesNotExist());
    }

    @Test
    void sameKeyWithDifferentWorkspaceSelectionConflicts() throws Exception {
        String tenant = "tenant-" + UUID.randomUUID();
        register(tenant, "ws-a", "storage-a");
        register(tenant, "ws-b", "storage-b");
        grant(tenant, "ws-a", "actor-a", true);
        grant(tenant, "ws-b", "actor-a", true);
        for (String[] selection : new String[][] {
                {"ws-a", "services/api", "202"},
                {"ws-a", "services/web", "409"},
                {"ws-b", "services/api", "409"}}) {
            mvc.perform(post("/v1/agents/sessions")
                            .header(TenantContextFilter.HEADER, tenant)
                            .header("Idempotency-Key", "same-key")
                            .principal(actor(tenant, "actor-a"))
                            .contentType(MediaType.APPLICATION_JSON)
                            .content("{\"agent_id\":\"qwen-code\",\"input\":[],"
                                    + "\"workspace\":{\"workspace_id\":\""
                                    + selection[0] + "\",\"cwd_relative\":\""
                                    + selection[1] + "\"}}"))
                    .andExpect(status().is(Integer.parseInt(selection[2])));
        }
    }

    @Test
    void readOnlyGrantCannotCreateBoundSession() throws Exception {
        String tenant = "tenant-" + UUID.randomUUID();
        register(tenant, "ws-a", "storage-a");
        grant(tenant, "ws-a", "reader", false);
        mvc.perform(post("/v1/agents/sessions")
                        .header(TenantContextFilter.HEADER, tenant)
                        .header("Idempotency-Key", "reader-create")
                        .principal(actor(tenant, "reader"))
                        .contentType(MediaType.APPLICATION_JSON)
                        .content("""
                                {"agent_id":"qwen-code","input":[],
                                 "workspace":{"workspace_id":"ws-a"}}
                                """))
                .andExpect(status().isForbidden())
                .andExpect(jsonPath("$.error.code")
                        .value("workspace_forbidden"));
        assertThat(jdbc.queryForObject("SELECT COUNT(*) FROM"
                        + " managed_agent_session WHERE tenant_id = ?",
                Integer.class, tenant)).isZero();
    }

F3 (design question, does not block): legacy and bound creation use separate Idempotency-Key namespaces

Live on the PR server, sending the same key with and without workspace creates two Sessions, and both responses say replay=false. This happens in either order. The controls behave correctly: within one mode, changed content gets 409 idempotency_conflict. A client or proxy that adds or drops workspace on retry will therefore get a duplicate Session silently. That goes against the test plan's "changing the request content should produce an idempotency conflict". Two options:

  • Reserve the key across both tables: bound creation also checks (tenant, CREATE_SESSION, key), and legacy creation checks the receipt table for (tenant, *, key).
  • Or state in the design doc that the two key spaces are separate.

mutation matrix and idempotency

Notes (do not block)

  • SSE with Accept: text/event-stream: any 404 becomes a 500, because ApiExceptionHandler cannot write JSON to an SSE response (HttpMediaTypeNotAcceptableException). This is pre-existing: main returns 500 for a nonexistent id too. It does mean a real EventSource client of a revoked actor gets 500 rather than the 404 the design describes; ?stream=true without that header gets 404. Mid-stream revocation closes the stream correctly, but logs 3 ERROR entries with full stack traces for an expected condition.
  • Tampered frozen configRef: it fails closed as designed (GET → 500). It also makes that actor's whole list return 500, so one bad row takes down the page. This is fine for W0b but worth knowing for operations.
  • I agree with triage on the default-Workspace branch (selection == null, managed_workspace_default): no HTTP route reaches it in this PR.
  • On triage note 3 (a SELECT per SSE event): M18, the per-event recheck on live delivery, is the only SSE guard mutant that survives, because the loop-top recheck covers it. Rechecking once before each batch would likely keep every tested behavior at one query per batch. I did not measure that.

Not covered

  • Production authenticated ingress, Worker mounts and bound execution. These are out of scope per the PR, and my rig's adapter trusts a header by design.
  • MariaDB locally. The CI MariaDB job is green. The index loss comes from a non-sargable predicate, which should behave the same way on MariaDB, but I only measured MySQL 8.4.7.
  • Windows and Linux locally. The CI Java jobs are green on ubuntu, macOS and windows.
中文说明

维护者验证:真实 MySQL + 真实服务 JAR,head cd76773f

结论:W0b 行为端到端正确(62/62 项实测)。建议合入前修复 F1。 F1 用实测数据回答了 triage 未能验证的 CAST 疑问(stage 2、stage 3)。这次改写确实让旧的未绑定路径用不上索引;FOR UPDATE 版本还会锁住所有租户的 command 行。一个 +29/−17 的候选补丁保留逐字节比较,同时恢复 main 的执行计划,已实测验证。F2(三个没有测试的守卫)和 F3(幂等键命名空间)问题较小,不阻塞合入。

做了什么

  • 环境: macOS arm64、JDK 21.0.12、Maven 3.9.16、全新 datadir 的 MySQL 8.4.7。对照臂为 origin/main 9e51e6ac,其 packages/sdk-java 与 PR base be1a0a1f 完全一致。
  • 测试套件: managed-agent-server 执行 mvn verify:69/69 通过,Checkstyle 无报错。-Pmysql-integration:默认 collation 的新库 6/6;utf8mb4_unicode_ci 库(feat(managed-agent): Spring control plane and dual-path WebShell #12692 F1 的那个 collation)也是 6/6。四张新表和 session 新增列均为 utf8mb4_bin。
  • 实机装置:
    • 真实 Spring Boot JAR,经 PropertiesLauncher 启动。
    • 一个 60 行的适配器 JAR,根据 X-Rig-Actor 设置 AuthenticatedTenantActor principal,扮演宿主侧已验证的入口;独立服务本身没有这一层。它按设计信任该 header,因此不能替代生产鉴权 E2E。
    • 一个计数用的假 Hosted Harness:记录每个请求,一律回 503。
    • Registry 行、授权和租户默认值由管理员通过 SQL 写入。

评审测试计划实测(图 1)

62 项全部通过:

  • 准入把 ./services/./api/ 规范化为 services/api,并保存七字段绑定和冻结的 configRef/policyRef。没有写入 Turn 行,Harness 没有收到请求。
  • generation=2、DRAINING、默认值变更之后重试,返回原 Session 和原绑定;改变 cwd 返回 409 idempotency_conflict。
  • actor 只取自 principal。通过 header 或 JSON body 声明 actor 得到 401;租户不匹配得到 403;Alice、ACME 大小写变体得到 404。
  • 10 种非法选择全部在写入任何行之前被拒绝。
  • 所有绑定 Turn 和生命周期操作:有读权限者得到 409,无读权限者得到 404。没有写入 command 或 Turn 行,Harness 没有收到请求。
  • 我直接在 MySQL 插入一条绑定的 ACCEPTED Turn,模拟已持久化的 Turn。调度器把它标为 FAILED workspace_unavailable,Harness 未被调用。
  • 撤销读权限后,所有读取路由返回 404,列表中也不再出现该 Session。已打开的 SSE 流在撤销后约 2 秒关闭。
  • 旧的未绑定路径仍然能到达 Harness。这是阳性对照,证明计数器有效。

F1(合入前修复):CAST(CONCAT(col,'!') AS BINARY(513)) 谓词在 MySQL 上用不了索引(图 2)

条件:10 万旧 Session(1000 个租户)、10 万 command 行、2 万条授权;各臂数据相同。

  • 执行计划:

    • listSessions 从 managed_agent_session_list_idx 上的 ref(100 行)变为全表扫描 99,375 行加 filesort。新增的授权 EXISTS 对每个绑定行都扫描全部 2 万条授权。
    • findCommand … FOR UPDATE 从 const 主键查找变为全表扫描 98,670 行。canRead 同样是全表扫描。
    • findSession/requireSessionForUpdate 不受影响,因为仍然能走 UNIQUE(session_id)。triage 对这点的判断是对的。
  • 锁: REPEATABLE READ 下,PR 对租户 t0500 执行 findCommand … FOR UPDATE,持有 101,650 个记录锁(main 只有 1 个)。此时租户 t0001 查询自己的记录、以及向 managed_agent_command 插入,都报 ERROR 1205 Lock wait timeout。这条语句在 insertTurnCommand、beginSessionMutation、completeSessionMutation 中都会执行,因此旧路径的 Turn 提交和所有生命周期修改会在全部租户之间串行化。

  • 真实 HTTP 延迟(中位数):

    路径 main PR 倍数
    列表 6.9 ms 88.9 ms 13×
    创建 5.8 ms 46.9 ms 8×
    16 个租户 × 10 次并行提交 Turn 39 ms(总耗时 0.7 s) 687 ms(总耗时 7.0 s) 17×
    含 300 个绑定 Session 的列表 — 687 ms —

    以上都随表大小线性增长,而 command 表只增不减。

  • CAST 的作用: 它不是无用的。MySQL 8.4 的 utf8mb4_bin 是 PAD SPACE,普通 = 下 'ws-a' = 'ws-a ' 为 TRUE,所以它为 Workspace ID 提供了精确比较。对租户 ID 和幂等键它没有实际作用,因为这两者的校验规则本来就不允许空格。

  • 候选修复: 在 7 个查询里,于 CAST 谓词前加一个可走索引的 col = ?,保留 CAST。CAST 接受的行一定也满足新增的 =,所以结果不变。

    • 已验证:单元/H2 69/69、MySQL IT 6/6、Checkstyle 无报错、实测 62/62。
    • 恢复了 main 的执行计划:列表走 ref list_idx,findCommand 走 const 主键,授权查询走 eq_ref/const 主键。
    • 锁测试恢复为 1 个锁,不再有跨租户等待。延迟:列表 7.6 ms、创建 5.1 ms、并行提交中位数 24 ms / 总耗时 0.4 s、绑定列表 9.5 ms。
  • 补丁见英文部分的折叠块。

F2(建议,只改测试):三个承重守卫没有测试(图 3)

我在 cd76773f 的副本上每次去掉一个守卫,跑 PR 自带的 66 个单元/H2 测试;这组运行不包含 MySQL IT。28 个变异体中 20 个被杀死。

  • 纵深防御(4 个存活): M02、M09、M11、M18。每个都还有上游守卫兜住。例如去掉服务层检查后,Store 的非空检查仍然返回 409。
  • 没有测试的完整性检查(1 个): M23,即重新推导冻结描述符。实测有效:被篡改的行会失败关闭。
  • 承重(3 个): 我为每个变异体构建了服务 JAR,在全新 MySQL 上运行。每个变异体只改变它对应的那一行;PR 代码在这三项上都正确。
    • M12: 撤销读权限后,WebShell 重试创建得到 202 {sessionId: <原 ID>, replayed: true},而不是 404。公开路由仍是 404,只是因为控制器会重新读取 Session。
    • M14: 同一个键把 ws-a/services/api 换成 ws-b/docs,得到 202 replay=true,并静默保留原绑定,而不是 409。
    • M24: 只读授权也能创建绑定 Session(202,而不是 403)。

三条候选测试(ManagedWorkspaceAdmissionTest +74 行)在 PR 代码上全部通过(66 → 69),并杀死 M12、M14、M24。代码见英文部分的折叠块。

F3(设计问题,不阻塞):旧创建和绑定创建使用不同的幂等键命名空间

在 PR 服务上实测:同一个键带和不带 workspace 各发一次,会创建两个 Session,两次响应都是 replay=false,先后顺序反过来也一样。对照组正常:同一模式下内容变化得到 409 idempotency_conflict。因此,客户端或代理在重试时加上或去掉 workspace,会静默得到重复的 Session,这与测试计划中「改变请求内容应得到幂等冲突」不符。两个方案:

  • 在两张表之间统一占用键:绑定创建也检查 (tenant, CREATE_SESSION, key),旧创建也检查回执表中的 (tenant, *, key)。
  • 或者在设计文档中写明两者是独立的键空间。

备注(不阻塞)

  • 带 Accept: text/event-stream 的 SSE: 任何 404 都会变成 500,因为 ApiExceptionHandler 无法向 SSE 响应写 JSON(HttpMediaTypeNotAcceptableException)。这是既有问题: main 对不存在的 ID 也返回 500。但这意味着被撤销权限的真实 EventSource 客户端拿到的是 500,而不是设计文档所说的 404;不带该 header、用 ?stream=true 时是 404。流进行中撤销权限时,流会正确关闭,但对这种预期内的情况会打 3 条带完整堆栈的 ERROR 日志。
  • 篡改冻结的 configRef: 会按设计失败关闭(GET → 500)。但它也会让该 actor 的整个列表返回 500,一行坏数据就让整页不可用。W0b 阶段可以接受,但运维上需要知道。
  • 我同意 triage 对默认 Workspace 分支(selection == null、managed_workspace_default)的判断:本 PR 中没有任何 HTTP 路由能到达它。
  • 关于 triage 第 3 点(每个 SSE 事件都执行一次 SELECT):M18(实时投递时的逐事件复查)是 SSE 守卫中唯一存活的变异体,因为循环顶部的复查已覆盖它。改为每批发送前复查一次,大概率能保留所有已测试的行为,同时每批只需一次查询。这一点我没有实测。

未覆盖

  • 生产鉴权入口、Worker 挂载、绑定执行。这些不在本 PR 范围内;我的装置中的适配器按设计信任 header。
  • 本地 MariaDB。CI 的 MariaDB 作业为绿。索引失效源于谓词无法走索引,在 MariaDB 上应该表现相同,但我只实测了 MySQL 8.4.7。
  • 本地 Windows 和 Linux。CI 的 Java 作业在 ubuntu、macOS、windows 上均为绿。

@wenshao

wenshao commented Sep 26, 2026

Copy link
Copy Markdown
Collaborator

Fixed in 02c2c82 and aa3ef35 (negative SSE cursor validation).

The three failing managed-agent CI tests are addressed: the Workspace identifier expectation now matches the frozen API contract; SSE revocation coverage verifies actual grant state and graceful completion; the batch projection fixture acquires its manual lease in the same transaction as admission so the background dispatcher cannot steal it first.

Correctness fixes also cover cross-shape creation idempotency (including concurrent requests and legacy receipt migration), indexed exact-match lookups, complete binding constraints, invalid Registry data classification, and SSE delivery through deletion. Session streams remain open across terminal Turn events and recheck bound grants before every event. English and Chinese designs are synchronized.

Verification / E2E report

  • Fresh-schema MariaDB 10.11 + Java 21 run: 85 unit/integration tests + 6 MariaDB IT cases passed, no failures/errors/skips; Checkstyle clean. After the final cursor-only correction, the full module rerun passes 86 unit/integration tests, with zero Checkstyle violations. Database code is unchanged by that correction.
  • Repository build, typecheck and bundle passed; build/typecheck were rerun after the final Java source correction.
  • Independent MariaDB probes: both sequential and forced-overlap bound/unbound creation orders yield one Session and one conflict; 10 simultaneous races pass. NULL numeric binding fields are rejected.
  • Actual JDBC SQL was captured and explained: primary-key, constant one-row access for command, Registry/ACL and Workspace receipt reads. A different tenant's locking lookup completed in 2 ms while the first transaction held its lock; the baseline timed out.
  • Six copied-module mutation controls turned the targeted regressions red when the relevant guards/parsing were broken. An additional SSE control failed all four delivery variants when turn.completed prematurely closed the stream; the final implementation passes all 11 SSE tests. The added negative-cursor test also failed before its validation guard was restored.
  • The affected MariaDB exact-case test passes independently on a fresh schema, without relying on test order.

Scope of evidence: HTTP behavior was exercised through MockMvc with trusted test principals. The final cursor fix was independently checked through both real controllers and async dispatch (400 invalid_event_cursor), with valid cursor controls. Full deployed authenticated Hosted Harness E2E is unavailable in this slice because the production ingress adapter and hosted session loop are not implemented; the report does not claim it ran. Database concurrency probes used real MariaDB.

Review handling: 29 existing threads are fixed or explicitly clarified and resolved. Eight nonblocking suggestions remain open with an individual rationale: R1-4 (catalog adapter refactor), R1-6 (error taxonomy), R1-22 (lazy digest optimization), R1-24 (fixture constants), R1-30 (broader cleanup-error preservation), R1-32 (ingress adapter diagnostics), R1-33 (shared-lock policy), and R1-38 (process launcher consolidation). These are deliberately deferred rather than represented as implemented.

GitHub CI for the new head is running; the local results above are independent of its pending result.

@wenshao

wenshao commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

Lease precision follow-up: 286bb8e.

The latest CI run exposed an additional lease-precision defect in the existing cross-process recovery test. This was reproduced rather than retried away: Connector/J sees MariaDB's 5.5.5 compatibility version and removes fractional seconds from Timestamp parameters. A nominal 1000 ms lease persisted about 800 ms early, so a commit 300 ms later failed even though the returned grant still advertised roughly 690 ms remaining.

The follow-up preserves fractional precision at all four lease-deadline writes (initial acquisition, reacquisition, renewal and takeover) by binding whole-second timestamps through JDBC and adding the preserved microseconds in SQL. This keeps the existing JDBC connection-timezone conversion. The database clock, lease duration, fencing and process-loss proof remain unchanged. The existing precision integration test now compares the persisted deadline with the returned grant for all four paths; it failed before the fix.

Final local verification: 86 unit/integration tests + 7 MariaDB IT cases passed on a fresh schema, with no failures/errors/skips and zero Checkstyle violations. The precision regression is parameterized over UTC and +08:00 connection timezones. Independent probes cover default and explicit connection timezones with UTC/Asia-Shanghai JVMs, including connection and JVM timezones that differ; all four paths retain the advertised deadline and delayed commits remain valid. Repository build/typecheck and independent code review were also rerun.

The six timezone combinations produced 42 exact deadline matches and 18 successful delayed commits. The SDK Java workflow for 286bb8e3de is now fully green, including all Java matrix jobs, the MariaDB job (86 + 7 managed-agent tests), and the real daemon E2E. General Qwen Code CI and the automated PR review are still running; no check is currently failing.

@wenshao

wenshao commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

Follow-up CI fix: 1fadbc6c4d makes the Linux unit-test commands honor filesystem permissions when the runner uses UID 0. Both workspace and script tests run with CAP_DAC_OVERRIDE and CAP_DAC_READ_SEARCH removed from the bounding and inheritable sets; ownership of the job checkout and isolated test home is normalized first.

The failed Node run had 11 cases where chmod-based read/delete/rename denial unexpectedly succeeded. All 11 passed under a normal local user, and a temporary permission-relaxing control reproduced all 11 CI failures. Local validation passed: 261 workflow tests, 18 execution-path checks of the actual workflow shell script with stubbed external commands, actionlint/ShellCheck, formatting, build, typecheck, and bundle.

Update: the Linux Node job passed on 1fadbc6c4d. All eight previously failing test files passed, and the script suite passed 2,627 tests across 96 files (12 conditional skips). The complete Qwen Code CI workflow, including Web Shell browser smoke, and the complete SDK Java workflow both finished successfully on the same head. Automatic PR review remains in progress.

Verification limitation: this attempt ran as the normal runner user (/home/github-runner), so the root-only capability branch was not executed. Its command wiring was reviewed and exercised with local stubs, but actual root-run setpriv validation remains unperformed.

wenshao added a commit to wenshao/qwen-code that referenced this pull request Sep 26, 2026
@wenshao

wenshao commented Sep 26, 2026

Copy link
Copy Markdown
Collaborator

Maintainer verification, round 2: head 1fadbc6c (packages/ identical to 286bb8e3)

Follow-up to round 1, same rig: real server JAR, MySQL 8.4.7, trusted-actor adapter, counting fake Hosted Harness.

Verdict: F1 and F3 are fixed and hold on real MySQL, including concurrency and an upgrade from a main-schema database. The SSE session.deleted regression flagged by bot finding R1-42 is also fixed; it was real, and my round 1 did not test deletion. Behavior is 68/68 live. From my side the only remaining item is one small test: M12, the replay-time read check, is still load-bearing and untested, and my round-1 test for it ports unchanged. Merge still needs the bot's CHANGES_REQUESTED review cleared, plus the still-pending Node CI for the new ci.yml change.

What changed since round 1

  • Fix commits (all in managed-agent-server): 7 commits covering:
    • indexable lookups (F1);
    • a single Idempotency-Key scope table with a V9 backfill (F3);
    • the W0a identifier rule at the API edge;
    • SSE changes: grant recheck with no per-event SELECT on legacy streams, graceful completion, delivery through deletion, and negative-cursor validation;
    • NOT NULL checks in V7;
    • Registry-row classification;
    • a lease-precision change in the shared ManagedSessionStore.
  • 1fadbc6c: touches only .github/workflows/ci.yml (git diff --quiet 286bb8e3 1fadbc6c -- packages docs is clean), so every measurement below applies to the current head.
  • Main: main has moved, but its only sdk-java change is docs plus one runtime-broker test, and git merge-tree is clean.

Suites

  • mvn verify: 86/86, Checkstyle clean.
  • MySQL IT: 7/7 on a fresh schema in each of 5 setups: JVM Asia/Shanghai or UTC × server CST or +00:00, plus one utf8mb4_unicode_ci database. The combinations exercise the lease-precision change; on MySQL, that change produces no regression.

Live walk-through: 68/68 (Fig. 1)

This is round 1's 62 checks plus new cases.

  • 'ws-a ' now gets 400 at the API edge (it was 404), as intended.
  • 'ws/a' gets 400.
  • New section 10 covers cross-shape keys:
    • unbound → bound: 409;
    • bound → unbound: 409;
    • WebShell unbound → public bound: 409;
    • the same key from a second actor still gets its own Session;
    • an identical retry replays.
  • The positive control shows the fake Harness is reached only by the legacy Turn.

round 2 live walk-through

Round-1 findings (Fig. 2)

Item Status on R2 Evidence
F1 index loss / table-wide locks ✅ fixed I captured the server's own SQL (general_log) and ran EXPLAIN on it: list uses a ref on list_idx, the grant subquery eq_ref PK, findCommand … FOR UPDATE and the create scope const PK. Lock test: 1 record lock, and another tenant's lookup/INSERT takes 0.02 s (R1: 101,650 locks, ERROR 1205). HTTP medians, main → R2: list 7.5 → 7.8 ms, create 5.3 → 5.8 ms, 16×10 parallel Turn submits 22.4 → 19.9 ms (0.4 s wall each). Bound list 8.8 ms (R1: 687 ms).
F3 Idempotency-Key namespaces ✅ fixed Live cross-shape 409s (section 10). Races: 30 rounds × 20 concurrent same-key requests (mixed, all-bound, all-unbound) gave exactly one Session per key in 30/30, with 0 5xx and 0 deadlocks; every loser of the other shape got 409. Real upgrade: the main JAR built V1–V6 with legacy Sessions, then the R2 JAR applied V7–V9. The backfilled scope made a bound create with a pre-upgrade key return 409, and a legacy retry replayed the same pre-upgrade Session.
F2 untested load-bearing guards ◑ 2 of 3 fixed M14 and M24 are now killed by the new tests. M12 is still untested (see below).
SSE revocation logging (round-1 note) ✅ fixed A mid-stream revocation now completes the stream about 2 s later with no ERROR log (R1: 3 ERROR entries with stack traces).
SSE Accept: text/event-stream → 500 (round-1 note) unchanged Pre-existing on main; not in scope.
SSE across deletion (bot R1-42) ✅ fixed Live A/B on a legacy stream: main delivers session.deleted and stays open. R1 closed after session.delete.requested without delivering session.deleted, with 6 ERROR lines. R2 delivers session.deleted, then completes, with 0 ERRORs. This held on both the public and WebShell routes.

round 1 findings re-measured

Mutation matrix: 35/38 killed (round 1: 20/28) (Fig. 3)

The matrix covers 38 single-guard mutants: round 1's 28 (5 re-targeted to the new code) plus 10 guards added by the fix commits. Every new guard is killed, including M38, which reverts findCommand to CAST-only; a test now catches that. Three mutants survive:

  • M12, the replay-time read check (load-bearing, no test). I built an R2 JAR without this check. On the real stack, a revoked actor retrying WebShell create gets 202 {sessionId: <original>, replayed: true}; R2 returns 404. The public route stays 404 only because its controller re-reads the Session. My round-1 test webShellRetryAfterRevocationReturnsNoSession ports unchanged: R2 goes 83 → 84 green, and the test kills M12.
  • M16, the loop-top SSE grant recheck (observable, no data leak). Without it, a revoked actor's idle bound stream stays open and keeps receiving :keepalive frames; it was still open 38 s later. R2 closes it at +2 s. The per-event checks still stop any event delivery, so no data leaks. A test is optional.
  • M11 (defense in depth). The store's non-empty check still returns 409.

round 2 mutation matrix

Test for M12, against the current test file (+28 lines)
    @Test
    void webShellRetryAfterRevocationReturnsNoSession() throws Exception {
        String tenant = "tenant-" + UUID.randomUUID();
        register(tenant, "ws-a", "storage-a");
        grant(tenant, "ws-a", "actor-a", true);
        String body = """
                {"idempotencyKey":"web-retry","agentId":"qwen-code",
                 "input":[],"workspace":{"workspaceId":"ws-a"}}
                """;
        mvc.perform(post("/api/agent/web-shell/v1/sessions/create")
                        .header(TenantContextFilter.HEADER, tenant)
                        .principal(actor(tenant, "actor-a"))
                        .contentType(MediaType.APPLICATION_JSON)
                        .content(body))
                .andExpect(status().isAccepted());
        jdbc.update("UPDATE managed_workspace_access SET can_read = FALSE"
                        + " WHERE tenant_id = ? AND workspace_id = ?"
                        + " AND actor_id = ?", tenant, "ws-a",
                "actor-a".getBytes(java.nio.charset.StandardCharsets.UTF_8));
        mvc.perform(post("/api/agent/web-shell/v1/sessions/create")
                        .header(TenantContextFilter.HEADER, tenant)
                        .principal(actor(tenant, "actor-a"))
                        .contentType(MediaType.APPLICATION_JSON)
                        .content(body))
                .andExpect(status().isNotFound())
                .andExpect(jsonPath("$.sessionId").doesNotExist());
    }

Notes (do not block)

  • V7 migration: ADD CONSTRAINT … CHECK cannot run INPLACE on MySQL 8.4 (ERROR 1845), so it copies the table and writes to managed_agent_session wait for it. Measured on 500,000 legacy Sessions: V7 takes 3.8 s and the V9 backfill 2.1 s. That is fine at this size; for much larger tables, plan the upgrade window.
  • Legacy SSE behavior change: legacy (unbound) streams now complete after session.deleted, where main kept them open. That seems right, but it is a change to already-shipped behavior, so it may deserve one line in the design notes.
  • ci.yml scope: 1fadbc6c is a repo-wide change to the Node test job, outside W0b's scope. Its chown stays inside the job workspace and the per-job HOME (runner.temp/qwen-ci-home). The Java jobs are all green on 1fadbc6c, including MariaDB; Test (ubuntu-latest, Node 22.x), the job this change affects, was still pending when I checked.
  • The tampered-descriptor case is unchanged apart from a better error message. I did not re-test it.

Not covered

  • Production authenticated ingress, Worker mounts and bound execution (out of scope per the PR).
  • MariaDB locally (the CI MariaDB job is green).
  • Linux/Windows locally.
  • The Linux capability behavior of the ci.yml change.
中文说明

维护者验证第 2 轮:head 1fadbc6c(packages/ 与 286bb8e3 完全相同)

接第 1 轮,装置相同:真实服务 JAR、MySQL 8.4.7、可信 actor 适配器、计数用的假 Hosted Harness。

结论:F1 和 F3 已修复,在真实 MySQL 上成立,包括并发场景和从 main 结构库升级的场景。 bot 的 R1-42 指出的 SSE session.deleted 回归也已修复;这个回归确实存在,我第 1 轮没有测删除。实测 68/68。我这边剩下的只有一个小测试:M12(重放时的读权限检查)仍然承重却没有测试,我第 1 轮写的测试可以原样移植。合入前还需要清掉 bot 的 CHANGES_REQUESTED 评审,并等新 ci.yml 改动对应的 Node CI 跑完(尚在运行)。

第 1 轮以来的变化

  • 修复提交(都在 managed-agent-server 内): 共 7 个,包括:
    • 让查询能走索引(F1);
    • 统一的幂等键作用域表,并在 V9 中回填(F3);
    • 在 API 入口校验 W0a 标识符规则;
    • SSE 相关修改:只复查授权、旧流不再每个事件执行一次 SELECT、正常结束流、删除后仍投递事件、校验负游标;
    • V7 增加 NOT NULL 检查;
    • Registry 行错误的分类;
    • 共享 ManagedSessionStore 的租约精度修改。
  • 1fadbc6c: 只改了 .github/workflows/ci.yml(git diff --quiet 286bb8e3 1fadbc6c -- packages docs 无差异),因此下面所有测量都适用于当前 head。
  • main: main 前进了,但它在 sdk-java 下只改了文档和一个 runtime-broker 测试,git merge-tree 无冲突。

测试套件

  • mvn verify:86/86,Checkstyle 无报错。
  • MySQL IT:5 种配置各用全新库,均为 7/7。配置是 JVM 时区 Asia/Shanghai 或 UTC × 服务端时区 CST 或 +00:00,外加一个 utf8mb4_unicode_ci 库。这些组合用来检验租约精度修改;在 MySQL 上它没有带来回归。

实测:68/68(图 1)

在第 1 轮 62 项的基础上增加了新用例。

  • 'ws-a ' 现在在 API 入口返回 400(之前是 404),符合预期。
  • 'ws/a' 返回 400。
  • 新增第 10 节测跨形态复用幂等键:
    • 先未绑定、再绑定:409;
    • 先绑定、再未绑定:409;
    • 先 WebShell 未绑定、再公开 API 绑定:409;
    • 另一个 actor 用同一个键,仍然得到自己的 Session;
    • 完全相同的重试会重放。
  • 阳性对照表明,只有旧路径的 Turn 会请求到假 Harness。

第 1 轮问题状态(图 2)

项 R2 状态 证据
F1 索引失效 / 全表锁 ✅ 已修复 我用 general_log 抓取服务实际执行的 SQL 并逐条 EXPLAIN:列表走 list_idx 的 ref,授权子查询走主键 eq_ref,findCommand … FOR UPDATE 和创建作用域都是主键 const。锁测试:1 个记录锁,另一个租户的查询/插入耗时 0.02 s(R1:101,650 个锁、ERROR 1205)。HTTP 中位数 main → R2:列表 7.5 → 7.8 ms,创建 5.3 → 5.8 ms,16×10 并行提交 Turn 22.4 → 19.9 ms(总耗时都是 0.4 s)。绑定列表 8.8 ms(R1:687 ms)。
F3 幂等键命名空间 ✅ 已修复 实测跨形态返回 409(第 10 节)。并发竞争: 30 轮 × 20 个同键并发请求(混合、全绑定、全未绑定),30/30 轮每个键恰好一个 Session,0 个 5xx,0 次死锁;另一形态的请求全部得到 409。真实升级: 先用 main JAR 建 V1–V6 并创建旧 Session,再用 R2 JAR 执行 V7–V9。回填后,用升级前的键做绑定创建返回 409,旧路径重试重放的是升级前那个 Session。
F2 承重守卫无测试 ◑ 3 个修了 2 个 M14、M24 现在都能被新测试杀死。M12 仍然没有测试(见下)。
SSE 撤销时的日志(第 1 轮备注) ✅ 已修复 流中途撤销权限后,约 2 秒流正常结束,没有 ERROR 日志(R1:3 条带堆栈的 ERROR)。
SSE Accept: text/event-stream → 500(第 1 轮备注) 未变 main 上既有的问题,不在本 PR 范围内。
SSE 跨删除投递(bot R1-42) ✅ 已修复 旧流实测 A/B:main 投递 session.deleted,之后流保持打开。R1 在 session.delete.requested 之后直接关闭,没有投递 session.deleted,并有 6 条 ERROR。R2 投递 session.deleted 后正常结束,0 条 ERROR。公开 API 和 WebShell 两条路由结果相同。

变异矩阵:38 个中杀死 35 个(第 1 轮:28 个中 20 个)(图 3)

共 38 个单守卫变异体:第 1 轮的 28 个(其中 5 个按新代码重新定位),加上修复提交新增守卫对应的 10 个。新增守卫全部能被测出,包括把 findCommand 退回纯 CAST 写法的 M38,现在也有测试能抓住。存活 3 个:

  • M12,重放时的读权限检查(承重,没有测试)。 我构建了去掉这项检查的 R2 JAR。在真实栈上,被撤销权限的 actor 重试 WebShell 创建会得到 202 {sessionId: <原 ID>, replayed: true};R2 返回 404。公开路由仍是 404,只是因为控制器会重新读取 Session。我第 1 轮写的测试 webShellRetryAfterRevocationReturnsNoSession 可以原样移植:R2 从 83 变为 84 全绿,并能杀死 M12。代码见英文部分的折叠块(+28 行)。
  • M16,SSE 循环顶部的授权复查(可观测,但不泄露数据)。 没有它时,被撤销权限的 actor 的空闲绑定流不会关闭,持续收到 :keepalive,38 秒后仍然打开;R2 在撤销后约 2 秒关闭。逐事件检查仍会阻止任何事件投递,因此不会泄露数据。测试可选。
  • M11(纵深防御)。 Store 层的非空检查仍然返回 409。

备注(不阻塞)

  • V7 迁移: MySQL 8.4 上 ADD CONSTRAINT … CHECK 不支持 INPLACE(ERROR 1845),因此会复制整张表,期间对 managed_agent_session 的写入要等待。在 50 万条旧 Session 上实测:V7 耗时 3.8 s,V9 回填 2.1 s。这个规模没有问题;表大得多时需要规划升级窗口。
  • 旧 SSE 行为变化: 旧的(未绑定)流现在在 session.deleted 之后结束,而 main 会保持打开。这样看起来是对的,但它改变了已上线的行为,也许值得在设计说明里写一行。
  • ci.yml 的范围: 1fadbc6c 是全仓库 Node 测试作业的改动,超出 W0b 范围。它的 chown 只作用于本作业的 workspace 和每个作业独立的 HOME(runner.temp/qwen-ci-home)。1fadbc6c 上的 Java 作业全部为绿,包括 MariaDB;受这项改动影响的 Test (ubuntu-latest, Node 22.x) 在我检查时仍在运行。
  • 篡改描述符的情况除错误信息更完整外没有变化,这次没有重测。

未覆盖

  • 生产鉴权入口、Worker 挂载和绑定执行(按 PR 说明不在范围内)。
  • 本地 MariaDB(CI 的 MariaDB 作业为绿)。
  • 本地 Linux/Windows。
  • ci.yml 改动在 Linux 上的 capability 行为。

agentService.lastSequence(tenantId, sessionId);
public SseEmitter publicStream(String tenantId, String actorId,
String sessionId, long afterSequence) {
SessionRecord session = agentService.requireReadableSession(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Unbound deleted Sessions now return 404 on events, items, transcript and SSE (legacy regression)

requireReadableSession goes through requireVisibleSession, which rejects status = 'DELETED'. On the merge base these routes never applied that filter:

  • publicEvents/webShellEvents did not check the session at all.
  • listPublicItems and transcript used store.requireSession, which returns DELETED rows.
  • Both stream routes used lastSequence → store.requireSession.

For an ordinary unbound Session this now means:

  • A client whose SSE connection dropped before the delete committed (a network blip, or the 30m stream-timeout) reconnects with its last sequence and gets 404 session_not_found. Before this change it received the terminal session.deleted event. It now can't tell "deleted" apart from "never existed".
  • Clients polling GET …/events?after=N, GET …/items, and WebShell transcript/query switch from 200 to 404 once the Session is deleted.

That contradicts "Unbound legacy behavior remains unchanged" in the PR body and the design doc. The R1-15 note that the unbound path "behaves exactly like the store.requireSession it replaced" doesn't hold for DELETED rows.

Suggested fix: for the event, item, transcript and stream-open paths, use a variant that does store.requireSession + requireReadGrant without the DELETED filter. Keep requireVisibleSession only for GET/list of the Session resource itself. Add a test where an unbound Session is deleted and a client then reconnects to its stream.

"Hosted Workspace execution is not available.");
}

private void requireCreationScope(String tenantId, String idempotencyKey,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Keys created by pre-V9 binaries during a rolling deploy skip the cross-shape conflict

V9 backfills managed_session_create_scope once, when the first upgraded instance runs Flyway. Instances still on the old binary keep inserting CREATE_SESSION receipts into managed_agent_command without a scope row. Later, insertWorkspaceSessionCommand finds no scope row for such a key K, inserts workspace_bound = TRUE, and admits a second Session under K. It should return idempotency_conflict (the invariant from the R1-2 fix).

The reverse also happens while old instances are still live: a bound create on a new instance, followed by an unbound create with K routed to an old instance (which has no scope check), produces two Sessions.

This is unlikely, because it needs the same key reused across shapes. But the gap is permanent for every key created during the rollout window.

Cheap fix: in the bound path, while holding the scope lock, also probe findCommand(tenantId, "CREATE_SESSION", idempotencyKey, true) and treat a hit as idempotency_conflict. That makes the legacy receipt table authoritative for legacy keys whether or not a scope row exists.

Comment thread .github/workflows/ci.yml
# Permission-denial fixtures must also work on root-run ECS lanes.
chown -R 0 "$GITHUB_WORKSPACE" "$HOME"
echo 'Dropping DAC override capabilities for root unit tests'
test_runner=(setpriv '--bounding-set=-dac_override,-dac_read_search' '--inh-caps=-dac_override,-dac_read_search')

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The DAC drop only covers this step

.github/scripts/run-release-workspace-tests.sh runs npm run test:release:workspaces on the same ecs-qwen-hk4-host self-hosted pool. It has no uid-0 check and no setpriv. On root release lanes, the permission-denial fixtures this change targets still run with CAP_DAC_OVERRIDE/CAP_DAC_READ_SEARCH, so they behave differently from PR CI. The release pipeline is the most expensive place to find that out.

Consider moving the chown + setpriv wrapper into a shared script, or adding it to the release script, so both lanes run tests with the same capability set.

Separately, this CI change doesn't depend on the W0b admission slice. A separate PR would let it be reviewed and reverted on its own.

AtomicBoolean leaseLost, AtomicBoolean submissionAttempted) {
SessionRecord session = store.requireSession(claimed.tenantId(),
claimed.sessionId());
if (session.workspace() != null) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This recovery guard can't run when the Hosted Harness is disabled

recoverExpiredTurns() returns immediately when !harness.isAvailable(), and the shipped default is QWEN_MANAGED_AGENT_HARNESS_ENABLED:false. In that configuration, an already-persisted bound Turn is never claimed, so it never reaches this fail(...). It stays ACCEPTED and keeps showing up as the Session's active_turn. The design doc says recovery "fails any already-persisted bound Turn before calling the Hosted Harness", which isn't true in this configuration.

If recovery is meant to fail closed whether or not the harness is available, fail bound dispatchable Turns before the isAvailable() early return. Otherwise, narrow the claim in the design doc (both languages).

long updatedAt, Long deletedAt, long version) {
long updatedAt, Long deletedAt, long version,
ContextBinding workspace) {
public SessionRecord(String tenantId, String sessionId,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The 13-arg convenience constructor quietly creates an unbound record

Every execution and visibility gate in this PR checks session.workspace() != null:

  • service: requireLegacyWorkspace, requireReadGrant
  • the four store guards
  • HarnessCoordinator
  • the embedded broker

This overload sets workspace to null, and only tests use it today. Future production code could rebuild a SessionRecord from an existing one through this overload, for example to copy it with a new status or version. That turns a bound Session into a legacy one and bypasses all of those gates with no compile error.

Consider removing the overload and passing the binding explicitly at the test call sites, or moving the helper into a test fixture.

ALTER TABLE managed_agent_session ADD COLUMN workspace_config_ref VARCHAR(512);
ALTER TABLE managed_agent_session ADD COLUMN workspace_policy_ref VARCHAR(512);

ALTER TABLE managed_agent_session ADD CONSTRAINT managed_workspace_binding_complete

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Adding this CHECK rebuilds managed_agent_session

The eight ADD COLUMN statements can run instantly, but ALTER TABLE … ADD CONSTRAINT … CHECK has to validate every existing row. MySQL can't do that instantly or in place; it copies the table. On a production-sized session table, this migration blocks writes to the busiest table for the length of the copy during deploy.

Either document the expected lock window in the design doc's migration notes, or combine the columns and the constraint into one ALTER TABLE and schedule the deploy with the rebuild in mind.

"sessionId", sessionId, "operation", operation));
}

private void requireLegacyWorkspace(String tenantId, String actorId,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Extra full-row read on every legacy mutation

requireLegacyWorkspace runs store.requireSession (a SELECT *) on every submit, cancel, rename, archive, unarchive and delete, including all unbound traffic. Then insertTurnCommand/insertCancelCommand/beginSessionMutation immediately re-read the same row FOR UPDATE and check workspace() != null again. The service-level read only adds one thing: choosing between 404 and 409 for bound Sessions.

A cheaper option: have the store guard throw a dedicated exception that carries the workspace id from the locked read, and map it to 404/409 in the service using canRead. That is one read per mutation on the legacy hot path instead of two, and the store stays the single owner of the gate.

String title, Map<String, Object> metadata,
List<InputBlock> blocks, WorkspaceSelection selection) {
validateIdempotencyKey(idempotencyKey);
if (actorId == null || actorId.isEmpty()) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit, per AGENTS.md: "No error handling for impossible scenarios."

  • Both controllers already call tenant.requireActorId() (twice each) before this method runs, and ManagedWorkspaceRegistry.resolveForCreation checks actorId == null || isEmpty() again. This third check can never trigger.
  • payloadDigest a few lines below is computed from input, which has just been confirmed empty, so it is always null. insertWorkspaceSessionCommand then gets a payloadDigest argument it can never use.

Removing the duplicate check and the dead payloadDigest leaves one place responsible for each rule.

+ " WHERE tenant_id = ?", tenant);
jdbc.update("DELETE FROM managed_workspace_access"
+ " WHERE tenant_id = ?", tenant);
jdbc.update("DELETE FROM managed_workspace_registry"

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Teardown doesn't clean up the scope rows

Every bound create in this test, and in @Order(6) (cleanup at L534), goes through requireCreationScope, which commits a managed_session_create_scope row. Neither finally block deletes from that table, so each run leaves scope rows in the shared MySQL schema. Tenants are UUID-randomized, so nothing collides today, but the cleanup no longer restores the schema the way it intends to.

Add DELETE FROM managed_session_create_scope WHERE tenant_id = ? to both blocks.

@wenshao

wenshao commented Sep 26, 2026

Copy link
Copy Markdown
Collaborator

Maintainer verification, round 3: Linux gap coverage, head 1fadbc6c

Verdict: merge-ready (round-3 scope). 21/21 scripted assertions passed, 0 unexpected failures. Verified head: 1fadbc6c4d2b4ac35bbe93781db84fd4e7967201.

Follow-up to round 2 at the same head. Round 2 (macOS arm64) left four items in Not covered: Linux/Windows locally, MariaDB locally, and the Linux capability behavior of the ci.yml change. This round runs on exactly the regime that was missing: Linux aarch64, UID 0 — the same regime as the root-run ecs-qwen lanes that 1fadbc6c targets. The live walk-through (68 checks, mutation matrix, SSE A/B) was not re-run: the head is byte-identical to round 2's, and those behaviors are platform-independent Java logic; instead I re-measured the load-bearing database claims (F1's fixed plans and lock footprint) independently on Linux, on both engines.

中文摘要

维护者验证第 3 轮:补齐 Linux 缺口,head 1fadbc6c

结论:merge-ready(本轮范围内)。21/21 项脚本断言全部通过,0 个意外失败。 验证的 head 为 1fadbc6c4d2b4ac35bbe93781db84fd4e7967201。

接第 2 轮(同一 head)。第 2 轮在 macOS 上留下了四个未覆盖项:本地 Linux/Windows、本地 MariaDB、ci.yml 改动的 Linux capability 行为。本轮正好在缺失的环境上运行:Linux aarch64、UID 0——与 1fadbc6c 针对的 root 运行 ecs-qwen 车道同一形态。第 2 轮的 68 项实测和变异矩阵没有重跑(head 完全相同,且那些是与平台无关的 Java 行为);我改为在 Linux 上独立复测了承重的数据库结论(F1 修复后的执行计划和锁占用),两种数据库引擎各测一遍。

本轮做了什么

  • Java 套件(图 1): runtime-broker 单元 326/326;managed-agent-server 单元 86/86,Checkstyle 0 违规——与第 2 轮 macOS 计数一致。MySQL 8.4 集成测试:broker 2/2、agent 7/7。MariaDB 11.4 集成测试(新覆盖):broker 2/2、agent 7/7。
  • F1 复测(图 2): 10 万 Session / 10 万 command / 2 万授权的种子数据上,逐字取自 head 源码的 SQL:findCommand … FOR UPDATE 在 MySQL 8.4 和 MariaDB 11.4 上都是主键 const(1 行);listSessions 走 list_idx 的 ref(100 行),授权子查询走主键 eq_ref/unique_subquery。锁测试:1 个记录锁(第 1 轮修复前为 101,650 个);另一租户的查询 0.4 ms、插入 9 ms,无 ERROR 1205。
  • ci.yml capability 修复(图 3、图 4): 这是 1fadbc6c 的主要内容,第 2 轮未覆盖。机制验证:root 直接读 chmod-000 文件成功(CAP_DAC_OVERRIDE 绕过);ci.yml 里的 setpriv 原样命令下读取返回 EACCES;chown -R 0 那一步确实必要(wrapper 下读别人拥有的文件是 EACCES,chown 后恢复)。真实套件 A/B:cleanup.test.ts 的 3 个权限夹具(只跳过 win32、不跳过 root)在普通 root 下 3 失败/31 通过——这正是 root 车道上原来的失败形态;在同一机器的 setpriv wrapper 下 34/34 全过。
  • 试合并: 与当前 origin/main 的 git merge-tree 干净无冲突;main 在 sdk-java 下只动了文档和一个一致性测试。

之前问题的状态

第 2 轮确认的所有修复(F1、F3、SSE 删除投递、撤销日志)在本轮未被削弱——F1 的执行计划与锁占用在 Linux 上两种引擎复测均成立(上图);其余修复针对平台无关的 Java 行为,head 与第 2 轮逐字节相同。第 2 轮遗留项不变:M12 测试仍缺失(不阻塞);bot 的 CHANGES_REQUESTED 评审仍待清除(流程项,非本轮实测范围)。

备注(不阻塞)

  • JdbcRuntimeBrokerMySqlIT 对非全新 schema 不可重跑:同一库上第二次运行失败(expected <[1]> but was <[2]>),删库重跑即恢复 2/2。CI 每次用全新库,无影响。
  • 本 PR base 的仓库还没有 package-lock.json(CI 用 corepack pnpm install --frozen-lockfile),本地验证需按此安装;main 已切换到 npm 锁文件,合并后自然消解。

未覆盖

  • 第 2 轮的完整 68 项实测 walk-through、变异矩阵、SSE A/B(理由同上,head 相同)。
  • Windows;生产鉴权入口、Worker 挂载、绑定执行(PR 声明的 W0c 范围外)。
  • Node 全套件(test:ci:workspaces 全量)——本轮只针对权限夹具文件做了 A/B;CI 的 Test (ubuntu-latest, Node 22.x) 在 head 上已绿。

Environment

  • Orange Pi (aarch64), Linux, UID 0, Oracle JDK 21 (build 21+35; round 2 used 21.0.12 on macOS), Maven 3.9.0, Node 24.14.0 + pnpm 11.24.0.
  • Docker mysql:8.4 and mariadb:11.4 containers on fresh schemas (this round's own instances).
  • Merge-base with origin/main is be1a0a1f — identical to baseRefOid, so no stale-base skew.

Suites on Linux (Fig. 1)

Suite MySQL 8.4 MariaDB 11.4
runtime-broker unit 326/326 (same arm)
managed-agent-server unit 86/86, Checkstyle 0 violations (same arm)
JdbcRuntimeBrokerMySqlIT 2/2 (fresh schema) 2/2
ManagedAgentMySqlIT 7/7 7/7

Unit counts match round 2's macOS counts exactly. The MariaDB 11.4 arm is new coverage — round 2 listed it in Not covered. V1–V9 migrations apply cleanly on both engines; all tables come out utf8mb4_bin.

java suites on linux

F1 re-measured on Linux, both engines (Fig. 2)

Seed: 100,000 sessions / 100,000 commands / 20,000 grants across 1,000 tenants (round 1's shape). SQL taken verbatim from ManagedAgentStore.java at head:

Check MySQL 8.4 MariaDB 11.4 Round-1 broken value
findCommand … FOR UPDATE plan const PRIMARY, 1 row const PRIMARY, 1 row full scan 98,670 rows
listSessions plan ref on list_idx, 100 rows ref on list_idx, 100 rows full scan 99,375 + filesort
grant EXISTS subplan eq_ref PRIMARY, 1 row unique_subquery PRIMARY, 1 row scan of all 20,000 grants
record locks held 1 1 (trx_rows_locked) 101,650 + ERROR 1205
cross-tenant lookup / INSERT 0.4 ms / 8.9 ms 0.5 ms / 4.5 ms lock-wait timeout

explain and locks

The ci.yml capability fix, verified in its target regime (Fig. 3, Fig. 4)

1fadbc6c is the head commit and was round 2's biggest Not covered item. This host is the regime it targets (root-run Linux lane).

  • Mechanism (Fig. 3): plain root reads a chmod-000 file (CapBnd: …fffff, all capabilities). The exact setpriv --bounding-set=-dac_override,-dac_read_search --inh-caps=… line from ci.yml drops the two bits (CapBnd: …ffff9) and the same read is EACCES. The chown -R 0 step is load-bearing: under the wrapper, a file owned by uid 65534 is EACCES until the chown restores owner-bit access — on a root lane whose checkout ran as another uid, the suite could not even read its own workspace without it.
  • Real-suite A/B (Fig. 4): packages/cli/src/utils/housekeeping/cleanup.test.ts has three chmod-000 permission fixtures that skip only on win32. Plain root: 3 failed / 31 passed — the failure shape the ECS root lane was hitting. Under the wrapper: 34/34 passed, same machine, same build.

capability mechanism

vitest capability A/B

Previous findings status

Item (round 2 state) Status this round Evidence
F1 fixed (index plans, locks) ✅ holds on Linux, both engines table + Fig. 2 above
F3 fixed (single key scope) not re-measured platform-independent Java behavior; head byte-identical to round 2
SSE across deletion fixed not re-measured same rationale
M12 test missing unchanged still open, still small and non-blocking per round 2
bot CHANGES_REQUESTED unchanged process item; a new review-pr run is pending on this head
ci.yml Linux capability behavior ✅ verified this round Fig. 3, Fig. 4
Node CI (Test (ubuntu-latest)) was pending ✅ now green on this head gh pr checks

Notes (do not block)

  • The broker MySQL IT is not re-run-safe against a non-fresh schema. A second run on the same database fails verifyBinding (expected <[1]> but was <[2]>); dropping the database restores 2/2. CI always uses a fresh schema, so this only bites local re-runs.
  • This base still installs with pnpm. package-lock.json does not exist at the PR's merge-base (CI uses corepack pnpm install --frozen-lockfile); main has since switched to the npm lockfile, so this resolves itself at merge.
  • Trial merge against current origin/main is clean (git merge-tree --write-tree); main's only sdk-java changes since the merge-base are docs plus one conformance test.

Not covered

  • Round 2's full live walk-through, mutation matrix, and SSE A/B — same head, platform-independent behavior; see the round-3 scope note at the top.
  • Windows locally; production authenticated ingress, Worker mounts, bound execution (W0c, out of scope per the PR).
  • The full Node suites (test:ci:workspaces / test:scripts end-to-end as root) — only the permission-fixture file was A/B'd; CI's Test (ubuntu-latest, Node 22.x) is green on this head.

Methodology

Worktree at 1fadbc6c (fetched from QwenLM/qwen-code), no source modifications. Java: mvn verify per module; ITs via -Pmysql-integration -Dmysql.url=… against this round's own MySQL 8.4 and MariaDB 11.4 containers, fresh schema per run (the one dirty-rerun failure is called out above). F1: migrations V1–V9 applied by hand, seeded via recursive CTEs, EXPLAIN of SQL copied verbatim from head source, lock counts via performance_schema.data_locks (MySQL) / information_schema.innodb_trx (MariaDB) with a two-connection holder/waiter rig. Capability: the exact setpriv line extracted from ci.yml, first against synthetic fixtures, then driving vitest on the real test file in both arms. Evidence captures via scripts/verify-capture.mjs; harnesses and raw logs live in tmp/pr12709-verify-20260926-104612/.

@wenshao

wenshao commented Sep 26, 2026

Copy link
Copy Markdown
Collaborator

@qwen-code /triage

@doudouOUC
doudouOUC added this pull request to the merge queue Sep 26, 2026
Merged via the queue into main with commit e944a23 Sep 26, 2026
327 of 336 checks passed
wenshao added a commit to doudouOUC/qwen-code that referenced this pull request Sep 26, 2026
…s-p0-p8

Brings in main through e944a23: QwenLM#12709 (workspace-bound Managed
sessions admitted without execution), QwenLM#12205 (MCP registration URL from
header discovery) and the v0.24.6 release.

The one conflict is EmbeddedRuntimeBrokerTest. It takes QwenLM#12709's test
that a workspace-bound Session cannot resolve the global runtime
workspace: this branch has no Hosted Workspace execution either, so the
new fail-closed rule in the embedded Broker's session resolver applies
here as it does on main. main's test that the embedded Broker answers
501 for prepare, start and resolve stays out, since this branch's Broker
serves those routes.

The rule merged into the branch's own EmbeddedRuntimeBroker without the
RuntimeBrokerException import that main's copy already had, so the
merge adds it.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants