Repository navigation
fix(compat): use nanosleep on POSIX in thread.sleep to actually suspend OS thread - #878
Conversation
e36ce0f to
2eb25a9
Compare
|
Cross-link for the NullHub-managed instance startup investigation: no new NullClaw code PR was needed beyond the existing downstream context here. NullClaw already contains the daemon nonblocking-accept fix; the remaining managed-instance failure in NullHub was caused by stale staged binaries. NullHub follow-up PR: nullclaw/nullhub#38. Provider/model-probe work remains in nullclaw/nullhub#34. |
|
Cross-link for the NullHub-managed instance startup investigation: No new NullClaw code PR was needed beyond the existing downstream context here. NullClaw NullHub follow-up PR: nullclaw/nullhub#38. |
…nd OS thread std.Io.sleep's cooperative yield does not actually suspend the OS thread in Threaded IO context. Use std.c.nanosleep directly on POSIX; keep std.Io.sleep on Windows where it routes through IOCP and works correctly. Regression test: sleep(0) early-return and sleep(1) smoke test.
2eb25a9 to
a77cbee
Compare
|
Hey @nullclaw/maintainers — this nanosleep fix has been sitting for a bit. It's a small one-file change that stops thread.sleep from busy-waiting on POSIX. Would love a review when you get a chance. Please. :-) |
|
I tried to reproduce the original sleep behavior issue in Docker, but I could not reproduce it. Environment:
I compared the old implementation: against the current Results: So, in this Linux Docker environment, the old I also tested the However, the So I can confirm that Could you share the exact OS, architecture, Zig version, libc/linking mode, and runtime context where the old |
|
Thanks. Here is the exact environment where I saw the original high-CPU behavior.
The 100%+ CPU symptom manifested for me on both macOS and the Radxa. I do not recall for sure whether I saw the same issue on the OrangePi. In my testing at the time, integrating this nanosleep change normalized CPU on the affected hosts. So the environments where I observed the problem were:
That is why I still believe the issue was real in production NullClaw runtimes even though it is not showing up in your Docker/Linux repro. |
|
One additional note: the current PR head now uses |
|
Thanks for the detailed environment notes. One thing still looks a bit strange to me: I checked Zig 0.16.0’s implementation in I also tried to reproduce the old behavior in Docker from Apple Silicon hosts, including MacBook M3/M4 setups, and I still cannot reproduce the high-CPU behavior there. In those environments the old That said, the current PR is small and the new implementation is explicit: on Linux/POSIX it calls So my current read is:
If the merged build keeps idle CPU normalized in the real agent/gateway process, that is good enough operationally, even if the exact underlying Zig/std.Io behavior remains unclear for me )) |
|
I didn't dig into the Zig code. It is strange that it uses the same posix method under the hood but there seemed to be divergent behavior. That's curious indeed. There was a bit of churn as we migrated to Zig 0.16.0 so it's possible something else was causing the issue. I'm only sure that after I introduced that change in my local environments the CPU utilization became normal again. |
Summary
std_compat.thread.sleep()to a real POSIXnanosleeppath instead ofstd.Io.sleep()'s cooperative yield understd.Io.ThreadedmainWhy
std_compat.thread.sleep()for backoff and pollingRefresh Status
mainon 2026-05-19a77cbeeValidation
zig test src/compat.zigzig build test --summary allis currently red on clean upstreammainin unrelated areas on this machine, so targeted validation is the meaningful signal for this PRNotes