What happened?
Since 9cb9dc86e8 (part of #12868, merged as 1f1209bc70), the Runtime Broker answers 200 with the state prepared for a provider execution whose dispatch the worker refused. Nothing will ever start that call, and the provider client waits for it without an end.
Steps, on the real chain (built TypeScript provider → Spring-embedded Runtime Broker on MySQL 8.4.7 → bundled worker), boot v2:
prepare a write_file and reserve it. Do not call preflight.
- Call
startExecution.
The worker answers execute with 409 managed_runtime_provider_operation_failed, "Managed tool preflight has not permitted execution.", as it should. What the caller sees:
|
50fb28301e, before the commit |
main a63157304a |
startExecution |
refused after 0.1 s, 409 runtime_broker_execution_unknown |
no answer; still waiting when the probe stopped after 20 s |
| Requests to the worker meanwhile |
1 execute |
1 execute, then 335 status in 20 s, about 17 a second |
| The Broker, asked directly |
start 409, read 409 |
start 200 prepared, read 200 prepared |
The same happens on boot v1 for a call that is started before it was approved. On 9cb9dc86e8 a probe of that case with no limit on its wait was stopped by hand after 4 min 26 s and 4,549 status requests, every one answered prepared.
Where it comes from: RuntimeBrokerHttpServer.observe now reconciles a provider record that is UNKNOWN, and observedExecutionEnvelope answers 200 for the states prepared, executing and cancel_requested. For a lost answer that is right, the worker says executing or settled. For a refused start the worker says prepared, and the Broker never dispatches a second time. The client's startExecution polls every 50 ms (EXECUTION_POLL_DELAY_MS) until the state is settled; only the provider's own lifetime signal ends that loop.
What did you expect to happen?
The caller of a refused start gets its refusal at once, as before the commit: 409 runtime_broker_execution_unknown, and the worker is not polled. A lost answer keeps settling from the result the worker kept, which is what the commit was written for.
Client information
Client Information
Measured on main a63157304a, and before the merge on 9cb9dc86e8, 760174073b and ebea694e4e. macOS 26 arm64, Node 22.23.2, JDK 21.0.12, MySQL 8.4.7. Provider client, Broker and worker are the ones built from that commit; a JVM HTTP proxy between Broker and worker keeps the ledger of requests.
Anything else we need to know?
Who is affected. No production code constructs BrokerManagedRuntimeProvider yet, so nothing that runs today is affected. A Hosted turn that starts a call it was not yet allowed to run would send exactly this sequence; before the commit that mistake failed closed in a tenth of a second, now it hangs the turn and loads the worker.
The caller has a way out since ebea694e4e. A cancellation reaches the worker, the record settles as not_started, and the Session is released. The start itself still does not return.
It reaches other checks. On main three probes of the earlier verification rounds fail because of it. A second start of a record that is UNKNOWN has no answer within 30 s, on boot v1 and on boot v2 (E-v1.2, E-v2.2). And a client that is still polling keeps sending status for as long as its provider lives, so a probe that counts the requests which reach the worker finds one it did not send (C-v1.5).
A candidate. One condition in observedExecutionEnvelope: a provider execution the worker still holds as prepared was never started, so the Broker keeps the answer it gave before. Measured with it on main a63157304a: the start is refused after 0.1 s, the worker gets one status request instead of 335, the probes Y1 to Y8 pass 21 of 21 (19 of 21 without it), the probes of the earlier rounds pass in full (76 of 76 and 12 of 12 where main has 75 and 10), the Broker suite passes with 514 tests, Checkstyle is clean. The test of the patch fails without the condition. Lost answers still settle and the cancellation of ebea694e4e still works.
Candidate patch (file)
diff --git a/packages/sdk-java/runtime-broker/src/main/java/com/alibaba/qwen/code/runtimebroker/RuntimeBrokerHttpServer.java b/packages/sdk-java/runtime-broker/src/main/java/com/alibaba/qwen/code/runtimebroker/RuntimeBrokerHttpServer.java
index 1ddaf1d258..772e1655c9 100644
--- a/packages/sdk-java/runtime-broker/src/main/java/com/alibaba/qwen/code/runtimebroker/RuntimeBrokerHttpServer.java
+++ b/packages/sdk-java/runtime-broker/src/main/java/com/alibaba/qwen/code/runtimebroker/RuntimeBrokerHttpServer.java
@@ -332,7 +332,12 @@ public final class RuntimeBrokerHttpServer implements AutoCloseable {
String runtimeSessionId, ExecutionReconciliation observation) {
ToolExecutionRecord record = observation.getRecord();
String state = observation.getRuntimeState();
- if (record.getState() == ToolExecutionRecord.State.UNKNOWN
+ // A provider call the worker still holds as prepared was never
+ // started, and nothing dispatches it a second time: it stays UNKNOWN
+ // for the caller, who would otherwise wait for it for ever.
+ boolean neverStarted = "prepared".equals(state)
+ && ProviderRuntimeProtocol.isReference(record.getReference());
+ if (record.getState() == ToolExecutionRecord.State.UNKNOWN && !neverStarted
&& ("prepared".equals(state) || "executing".equals(state)
|| "cancel_requested".equals(state))) {
Map<String, Object> response = envelope(harnessSessionId, runtimeSessionId,
diff --git a/packages/sdk-java/runtime-broker/src/test/java/com/alibaba/qwen/code/runtimebroker/RuntimeBrokerHttpServerTest.java b/packages/sdk-java/runtime-broker/src/test/java/com/alibaba/qwen/code/runtimebroker/RuntimeBrokerHttpServerTest.java
index 2be20af7c3..53464b377b 100644
--- a/packages/sdk-java/runtime-broker/src/test/java/com/alibaba/qwen/code/runtimebroker/RuntimeBrokerHttpServerTest.java
+++ b/packages/sdk-java/runtime-broker/src/test/java/com/alibaba/qwen/code/runtimebroker/RuntimeBrokerHttpServerTest.java
@@ -226,6 +226,34 @@ class RuntimeBrokerHttpServerTest {
}
}
+ @Test
+ void providerStartTheWorkerNeverBeganStaysUnknownForTheCaller() throws Exception {
+ String runtime = "550e8400-e29b-41d4-a716-446655440303";
+ try (Fixture fixture = new Fixture()) {
+ fixture.service.acquire("harness", runtime, "bootstrap").toCompletableFuture().join();
+ Map<String, Object> reference = Map.of("sessionId", runtime, "promptId", "turn",
+ "callId", "worker-call", "capabilityDigest", "a".repeat(64),
+ "policyRevision", "policy", "invocationId", "invocation", "argsDigest", "b".repeat(64));
+ String id = fixture.service.prepareExecution("harness", runtime, "provider", reference)
+ .toCompletableFuture().join().getExecutionCallId();
+ Map<String, Object> start = Map.of("protocolVersion", 1, "requestId", "start",
+ "harnessSessionId", "harness", "runtimeSessionId", runtime);
+ // The worker refused the execute: it still holds the call as prepared.
+ fixture.transport.runtimeStatus = Map.of("state", "prepared");
+ for (int attempt = 0; attempt < 2; attempt++) {
+ HttpResponse<String> started = fixture.post("/executions/" + id + ":start", start);
+ assertEquals(409, started.statusCode(), started.body());
+ assertEquals("runtime_broker_execution_unknown",
+ JSON.parseObject(started.body()).getString("code"));
+ }
+ HttpRequest read = HttpRequest.newBuilder(fixture.uri("/executions/" + id
+ + "?requestId=read&harnessSessionId=harness&runtimeSessionId=" + runtime))
+ .header("Authorization", "Bearer secret").GET().build();
+ assertEquals(409, fixture.client.send(read, HttpResponse.BodyHandlers.ofString()).statusCode());
+ assertEquals(1, fixture.transport.executions.get());
+ }
+ }
+
@Test
void providerObservesTheOriginalExecutionAfterResponseLoss() throws Exception {
// Provider references carry a UUID Runtime Session id.
The other place to close it would be the client: stop waiting when a call that was started reads prepared. That was not measured.
Evidence. Round 7 of the real-stack verification of #12868: #12868 (comment) (section R7-1 and figure 6). Logs: on main, on main with the candidate, before the commit, the long run, the earlier probes on main (contract, faults) and with the candidate (contract, faults). The probe is group Y8 of s23-merge-r7.mjs.
中文说明
发生了什么?
从 9cb9dc86e8 起(属于 #12868,已合入为 1f1209bc70),对于 worker 拒绝派发的 provider 执行,Runtime Broker 会以 200 加状态 prepared 作答。这个调用永远不会被启动,而 provider 客户端会无止境地等它。
在真实链路上(构建出来的 TypeScript provider → Spring 内嵌 Runtime Broker,MySQL 8.4.7 → 打包 worker),boot v2:
prepare 一个 write_file 并预留,不调用 preflight。
- 调用
startExecution。
worker 以 409 managed_runtime_provider_operation_failed(“Managed tool preflight has not permitted execution.”)拒绝 execute,这是正确的。调用方看到的是:
|
50fb28301e(该提交之前) |
main a63157304a |
startExecution |
0.1 秒后被拒绝,409 runtime_broker_execution_unknown |
没有应答;探针 20 秒后停止时仍在等待 |
| 这期间发给 worker 的请求 |
1 次 execute |
1 次 execute,随后 20 秒内 335 次 status,约每秒 17 次 |
| 直接问 Broker |
start 409,read 409 |
start 200 prepared,read 200 prepared |
boot v1 上,调用在获批之前就被 start 时也是一样。在 9cb9dc86e8 上,针对这种情形、不设等待上限的探针在 4 分 26 秒、4,549 次 status 请求之后被手动停止,每一次的回答都是 prepared。
原因:RuntimeBrokerHttpServer.observe 现在会对处于 UNKNOWN 的 provider 记录做对账,observedExecutionEnvelope 对 prepared、executing、cancel_requested 三种状态都返回 200。对于应答丢失的情形这是对的,worker 会说 executing 或 settled。但对于被拒绝的 start,worker 说的是 prepared,而 Broker 从不二次派发。客户端的 startExecution 每 50 毫秒轮询一次(EXECUTION_POLL_DELAY_MS),直到状态变成 settled;只有 provider 自己的生命周期信号能结束这个循环。
期望的行为
被拒绝的 start 应该像该提交之前那样立刻得到拒绝:409 runtime_broker_execution_unknown,并且不去轮询 worker。应答丢失的情形继续按 worker 保留的结果结算,这正是那个提交要解决的问题。
环境
在 main a63157304a 上实测;合入之前在 9cb9dc86e8、760174073b、ebea694e4e 上也实测过。macOS 26 arm64,Node 22.23.2,JDK 21.0.12,MySQL 8.4.7。provider 客户端、Broker 和 worker 都由对应提交构建;Broker 与 worker 之间挂着 JVM HTTP 代理,用来记录请求账本。
其他信息
影响范围。 目前没有任何生产代码构造 BrokerManagedRuntimeProvider,所以今天在跑的东西不受影响。如果 Hosted 的某个回合在调用尚未获准时就 start,发出的正是这个请求序列;在该提交之前这种错误会在十分之一秒内以失败告终,现在它会挂住整个回合,并持续给 worker 加压。
从 ebea694e4e 起调用方有了出路。 cancel 能到达 worker,记录结算为 not_started,Session 可以释放。但 start 本身仍然不返回。
它会波及别的检查。 在 main 上,此前各轮验证里有三个探针因它而失败。对 UNKNOWN 记录的第二次 start 在 30 秒内没有应答,boot v1 和 boot v2 都是如此(E-v1.2、E-v2.2)。另外,仍在轮询的客户端只要 provider 还活着就会不断发送 status,所以统计到达 worker 的请求数的探针会数到一条不是它自己发的请求(C-v1.5)。
候选修复。 在 observedExecutionEnvelope 里加一个条件:worker 仍然标记为 prepared 的 provider 执行其实从未启动,因此 Broker 保持原来的应答。在 main a63157304a 上加上它之后的实测:start 在 0.1 秒后被拒绝,worker 收到 1 次 status 请求而不是 335 次,探针 Y1 到 Y8 为 21/21(不加时为 19/21),此前各轮的探针全部通过(76/76 和 12/12,main 上是 75 和 10),Broker 套件 514 个测试通过,Checkstyle 干净。去掉这个条件,补丁自带的测试就会失败。应答丢失的情形仍然能结算,ebea694e4e 的取消也仍然有效。补丁见上面英文部分的折叠块。
另一个可以关闭它的位置是客户端:已经 start 的调用读到 prepared 时停止等待。这个方案没有实测。
证据。 #12868 真实环境验证第七轮:#12868 (comment) (R7-1 一节和图 6)。日志与探针的链接见上面英文部分。
What happened?
Since
9cb9dc86e8(part of #12868, merged as1f1209bc70), the Runtime Broker answers200with the statepreparedfor a provider execution whose dispatch the worker refused. Nothing will ever start that call, and the provider client waits for it without an end.Steps, on the real chain (built TypeScript provider → Spring-embedded Runtime Broker on MySQL 8.4.7 → bundled worker), boot v2:
prepareawrite_fileand reserve it. Do not callpreflight.startExecution.The worker answers
executewith409 managed_runtime_provider_operation_failed, "Managed tool preflight has not permitted execution.", as it should. What the caller sees:50fb28301e, before the commitmaina63157304astartExecution409 runtime_broker_execution_unknownexecuteexecute, then 335statusin 20 s, about 17 a second409, read409200 prepared, read200 preparedThe same happens on boot v1 for a call that is started before it was approved. On
9cb9dc86e8a probe of that case with no limit on its wait was stopped by hand after 4 min 26 s and 4,549statusrequests, every one answeredprepared.Where it comes from:
RuntimeBrokerHttpServer.observenow reconciles a provider record that isUNKNOWN, andobservedExecutionEnvelopeanswers200for the statesprepared,executingandcancel_requested. For a lost answer that is right, the worker saysexecutingorsettled. For a refused start the worker saysprepared, and the Broker never dispatches a second time. The client'sstartExecutionpolls every 50 ms (EXECUTION_POLL_DELAY_MS) until the state issettled; only the provider's own lifetime signal ends that loop.What did you expect to happen?
The caller of a refused start gets its refusal at once, as before the commit:
409 runtime_broker_execution_unknown, and the worker is not polled. A lost answer keeps settling from the result the worker kept, which is what the commit was written for.Client information
Client Information
Measured on
maina63157304a, and before the merge on9cb9dc86e8,760174073bandebea694e4e. macOS 26 arm64, Node 22.23.2, JDK 21.0.12, MySQL 8.4.7. Provider client, Broker and worker are the ones built from that commit; a JVM HTTP proxy between Broker and worker keeps the ledger of requests.Anything else we need to know?
Who is affected. No production code constructs
BrokerManagedRuntimeProvideryet, so nothing that runs today is affected. A Hosted turn that starts a call it was not yet allowed to run would send exactly this sequence; before the commit that mistake failed closed in a tenth of a second, now it hangs the turn and loads the worker.The caller has a way out since
ebea694e4e. A cancellation reaches the worker, the record settles asnot_started, and the Session is released. The start itself still does not return.It reaches other checks. On
mainthree probes of the earlier verification rounds fail because of it. A second start of a record that isUNKNOWNhas no answer within 30 s, on boot v1 and on boot v2 (E-v1.2, E-v2.2). And a client that is still polling keeps sendingstatusfor as long as its provider lives, so a probe that counts the requests which reach the worker finds one it did not send (C-v1.5).A candidate. One condition in
observedExecutionEnvelope: a provider execution the worker still holds aspreparedwas never started, so the Broker keeps the answer it gave before. Measured with it onmaina63157304a: the start is refused after 0.1 s, the worker gets onestatusrequest instead of 335, the probes Y1 to Y8 pass 21 of 21 (19 of 21 without it), the probes of the earlier rounds pass in full (76 of 76 and 12 of 12 wheremainhas 75 and 10), the Broker suite passes with 514 tests, Checkstyle is clean. The test of the patch fails without the condition. Lost answers still settle and the cancellation ofebea694e4estill works.Candidate patch (file)
The other place to close it would be the client: stop waiting when a call that was started reads
prepared. That was not measured.Evidence. Round 7 of the real-stack verification of #12868: #12868 (comment) (section R7-1 and figure 6). Logs: on
main, onmainwith the candidate, before the commit, the long run, the earlier probes onmain(contract, faults) and with the candidate (contract, faults). The probe is group Y8 ofs23-merge-r7.mjs.中文说明
发生了什么?
从
9cb9dc86e8起(属于 #12868,已合入为1f1209bc70),对于 worker 拒绝派发的 provider 执行,Runtime Broker 会以200加状态prepared作答。这个调用永远不会被启动,而 provider 客户端会无止境地等它。在真实链路上(构建出来的 TypeScript provider → Spring 内嵌 Runtime Broker,MySQL 8.4.7 → 打包 worker),boot v2:
prepare一个write_file并预留,不调用preflight。startExecution。worker 以
409 managed_runtime_provider_operation_failed(“Managed tool preflight has not permitted execution.”)拒绝execute,这是正确的。调用方看到的是:50fb28301e(该提交之前)maina63157304astartExecution409 runtime_broker_execution_unknownexecuteexecute,随后 20 秒内 335 次status,约每秒 17 次409,read409200 prepared,read200 preparedboot v1 上,调用在获批之前就被 start 时也是一样。在
9cb9dc86e8上,针对这种情形、不设等待上限的探针在 4 分 26 秒、4,549 次status请求之后被手动停止,每一次的回答都是prepared。原因:
RuntimeBrokerHttpServer.observe现在会对处于UNKNOWN的 provider 记录做对账,observedExecutionEnvelope对prepared、executing、cancel_requested三种状态都返回200。对于应答丢失的情形这是对的,worker 会说executing或settled。但对于被拒绝的 start,worker 说的是prepared,而 Broker 从不二次派发。客户端的startExecution每 50 毫秒轮询一次(EXECUTION_POLL_DELAY_MS),直到状态变成settled;只有 provider 自己的生命周期信号能结束这个循环。期望的行为
被拒绝的 start 应该像该提交之前那样立刻得到拒绝:
409 runtime_broker_execution_unknown,并且不去轮询 worker。应答丢失的情形继续按 worker 保留的结果结算,这正是那个提交要解决的问题。环境
在
maina63157304a上实测;合入之前在9cb9dc86e8、760174073b、ebea694e4e上也实测过。macOS 26 arm64,Node 22.23.2,JDK 21.0.12,MySQL 8.4.7。provider 客户端、Broker 和 worker 都由对应提交构建;Broker 与 worker 之间挂着 JVM HTTP 代理,用来记录请求账本。其他信息
影响范围。 目前没有任何生产代码构造
BrokerManagedRuntimeProvider,所以今天在跑的东西不受影响。如果 Hosted 的某个回合在调用尚未获准时就 start,发出的正是这个请求序列;在该提交之前这种错误会在十分之一秒内以失败告终,现在它会挂住整个回合,并持续给 worker 加压。从
ebea694e4e起调用方有了出路。 cancel 能到达 worker,记录结算为not_started,Session 可以释放。但 start 本身仍然不返回。它会波及别的检查。 在
main上,此前各轮验证里有三个探针因它而失败。对UNKNOWN记录的第二次 start 在 30 秒内没有应答,boot v1 和 boot v2 都是如此(E-v1.2、E-v2.2)。另外,仍在轮询的客户端只要 provider 还活着就会不断发送status,所以统计到达 worker 的请求数的探针会数到一条不是它自己发的请求(C-v1.5)。候选修复。 在
observedExecutionEnvelope里加一个条件:worker 仍然标记为prepared的 provider 执行其实从未启动,因此 Broker 保持原来的应答。在maina63157304a上加上它之后的实测:start 在 0.1 秒后被拒绝,worker 收到 1 次status请求而不是 335 次,探针 Y1 到 Y8 为 21/21(不加时为 19/21),此前各轮的探针全部通过(76/76 和 12/12,main上是 75 和 10),Broker 套件 514 个测试通过,Checkstyle 干净。去掉这个条件,补丁自带的测试就会失败。应答丢失的情形仍然能结算,ebea694e4e的取消也仍然有效。补丁见上面英文部分的折叠块。另一个可以关闭它的位置是客户端:已经 start 的调用读到
prepared时停止等待。这个方案没有实测。证据。 #12868 真实环境验证第七轮:#12868 (comment) (R7-1 一节和图 6)。日志与探针的链接见上面英文部分。