You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 56c2096
Browse filesBrowse the repository at this point in the historyBrowse files
wenshao
committed
fix(runtime-broker): repair the Windows leg and close the release grace window
Round 2 broke windows-latest / Java 21 (612 tests, 2 failures, with ubuntu
and macOS green). Both are fixed here without waiving the platform:
- The exit-hook test had moved the decision of when the forked harness
exits from a sleep inside it to harness.destroy() in the test. destroy()
is a SIGTERM only on POSIX; on Windows it is TerminateProcess, which runs
no shutdown hooks, so the worker legitimately outlived the harness. The
harness now polls a sentinel file and exits itself with System.exit(0)
once the test has observed the worker: the window stays test-controlled
and the hooks run on every platform.
- The provision-race test widens its race window with a --ignore-term
worker, which cannot wedge anything on Windows, so close() finishes in
milliseconds and the racing loop can miss the terminated guard. The guard
assertion now runs only where destroy() is a signal; the
platform-independent half - no lingering child processes - still runs
everywhere. That code is unchanged from the previous Windows-green
commit, so this was a timing flake rather than a new defect.
Also close a real orphan window in the release path. stop() removed the
worker from owned and then escalated to destroyForcibly on a daemon thread,
so for the whole five-second grace the process was tracked by neither set.
close() was covered, because shutting the executor down interrupts the
escalation and its handler destroys forcibly, but a bare JVM exit was not:
the daemon dies with the JVM and the hook's snapshot could not see the
worker. A worker that ignores SIGTERM, released and then hit by a Broker
SIGTERM inside five seconds, survived. The worker now moves from owned to
starting under the same lifecycle lock terminateAll() snapshots with, which
also removes the microsecond gap between those two set operations, and
leaves it when the escalation finishes. A forked test releases a wedged
worker and exits the harness JVM inside the grace window; without the
tracking it fails with that worker still alive.
Smaller corrections from the same audit round: the cross-process stress
rises to 600 rounds so its raced arm keeps the 200 rounds the odd/even
split had halved, and its comments now say what the win counters prove -
both arms ran - instead of claiming a contradictory interleaving was
reached; the v3 backoff test sweeps attempts 0-1000, pinning the inner
shift clamp (an unclamped shift wraps at attempt 57 and schedules zero and
negative delays, returning the polling to the rate the backoff exists to
remove); and the design doc's description of the pre-fix LOST reclaim is
corrected in both languages, since at the merge base it ran a single
bounded pass per phase and stranded larger generations rather than looping
without a budget. A test name and two assertion messages that still carried
rationales round 2 retracted in production are corrected, and the
exit-hook test's worker cleanup moves into its finally block so a failed
assertion cannot leave a spinning node process on the runner.
Declined: a production seam inside the release transaction to make the
row-lock pin deterministic. The raced arm already detects the missing
FOR UPDATE in three runs out of three, with three to five contradictions
each, and the transition's in-transaction re-check is pinned
deterministically by the delegating-repository release test.
Suites green: runtime-broker 631 tests / 0 failures / 2 skipped with
checkstyle and spotbugs clean, managed-agent-server fix-adjacent 19/19
against the reinstalled jar.
A code audit of the Runtime Broker found three high-risk defects. First, the session release decision and the no-active-execution check ran in separate transactions guarded only by a process-local lock: with two Broker processes sharing one database, an execution could be admitted while its session was marked `RELEASED`. Second, every lease renewal ran on one scheduled thread shared with retries, deadline fences, and polling, all with synchronous JDBC: a 1–2 s storage stall queued renewals past their lease and fenced healthy bindings. Third, the broker's HTTP face served one global Bearer token over plaintext HTTP and accepted non-loopback listen addresses.
10
10
11
-
Two cheaper medium findings are fixed in the same change. A released worker that ignored SIGTERM was never destroyed forcibly, and a JVM exit during the ready handshake could strand one. A LOST reclaim also repeated its bounded 100-row recovery passes with no budget, so a generation larger than one pass could spin.
11
+
Two cheaper medium findings are fixed in the same change. A released worker that ignored SIGTERM was never destroyed forcibly, and a JVM exit during the ready handshake could strand one. A LOST reclaim also ran a single bounded 100-row recovery pass per phase, so a generation larger than one pass stayed LOST and answered `runtime_broker_runtime_lost` on every later attempt.
Copy file name to clipboardExpand all lines: packages/sdk-java/runtime-broker/src/main/java/com/alibaba/qwen/code/runtimebroker/LocalProcessRuntimeProvisioner.java
+20-2Lines changed: 20 additions & 2 deletions
Original file line number
Diff line number
Diff line change
@@ -348,14 +348,32 @@ public boolean isUsable(RuntimeLease lease) {
Copy file name to clipboardExpand all lines: packages/sdk-java/runtime-broker/src/test/java/com/alibaba/qwen/code/runtimebroker/Issue13183AdversarialTest.java
0 commit comments