Skip to content

fix(compat): use nanosleep on POSIX in thread.sleep to actually suspend OS thread - #878

Merged
DonPrus merged 4 commits into
nullclaw:mainfrom
vernonstinebaker:fix/compat-thread-sleep
May 28, 2026
Merged

DonPrus merged 4 commits into
nullclaw:mainfrom
vernonstinebaker:fix/compat-thread-sleep

Conversation

@vernonstinebaker

@vernonstinebaker vernonstinebaker commented Apr 30, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • switch std_compat.thread.sleep() to a real POSIX nanosleep path instead of std.Io.sleep()'s cooperative yield under std.Io.Threaded
  • keep the existing scheduler-backed path for Windows/WASI
  • keep compat tests/build wiring working on current main

Why

  • NullClaw's managed gateway and channel loops rely on std_compat.thread.sleep() for backoff and polling
  • on POSIX with Zig 0.16, the current implementation still does not reliably suspend the OS thread, so this fix remains relevant

Refresh Status

  • rebased/refreshed onto current main on 2026-05-19
  • current PR head: a77cbee
  • no functional changes beyond the original fix; this is a branch refresh only

Validation

  • zig test src/compat.zig
  • full zig build test --summary all is currently red on clean upstream main in unrelated areas on this machine, so targeted validation is the meaningful signal for this PR

Notes

  • this remains separate from the already-merged accept-loop/backoff fixes
  • no other upstream PR appears to cover this POSIX sleep behavior change

@vernonstinebaker
vernonstinebaker force-pushed the fix/compat-thread-sleep branch from e36ce0f to 2eb25a9 Compare May 2, 2026 07:18
@vernonstinebaker

Copy link
Copy Markdown
Contributor Author

Cross-link for the NullHub-managed instance startup investigation: no new NullClaw code PR was needed beyond the existing downstream context here. NullClaw already contains the daemon nonblocking-accept fix; the remaining managed-instance failure in NullHub was caused by stale staged binaries. NullHub follow-up PR: nullclaw/nullhub#38. Provider/model-probe work remains in nullclaw/nullhub#34.

@vernonstinebaker

Copy link
Copy Markdown
Contributor Author

Cross-link for the NullHub-managed instance startup investigation:

No new NullClaw code PR was needed beyond the existing downstream context here. NullClaw main already contains the daemon nonblocking-accept fix; the remaining managed-instance failure in NullHub was caused by stale staged dev-local binaries.

NullHub follow-up PR: nullclaw/nullhub#38.
Provider/model-probe work remains in nullclaw/nullhub#34.

…nd OS thread

std.Io.sleep's cooperative yield does not actually suspend the OS thread
in Threaded IO context. Use std.c.nanosleep directly on POSIX; keep
std.Io.sleep on Windows where it routes through IOCP and works correctly.

Regression test: sleep(0) early-return and sleep(1) smoke test.
@vernonstinebaker
vernonstinebaker force-pushed the fix/compat-thread-sleep branch from 2eb25a9 to a77cbee Compare May 19, 2026 05:22
@vernonstinebaker

Copy link
Copy Markdown
Contributor Author

Hey @nullclaw/maintainers — this nanosleep fix has been sitting for a bit. It's a small one-file change that stops thread.sleep from busy-waiting on POSIX. Would love a review when you get a chance. Please. :-)

@DonPrus

DonPrus commented May 27, 2026

Copy link
Copy Markdown
Contributor

I tried to reproduce the original sleep behavior issue in Docker, but I could not reproduce it.

Environment:

  • Docker server: linux/arm64
  • Container base image: alpine:3.20
  • Kernel: Linux 6.12.76-linuxkit aarch64
  • Zig: 0.16.0
  • libc: not linked (builtin.link_libc=false)
  • test mode: false (builtin.is_test=false)

I compared the old implementation:

std.Io.sleep(compat.io(), .fromNanoseconds(@intCast(nanoseconds)), .awake) catch {};

against the current compat.thread.sleep(...) implementation.

Results:

old_process_init_elapsed_ns=204132583
old_explicit_threaded_elapsed_ns=205466584
new_current_elapsed_ns=205023166

So, in this Linux Docker environment, the old std.Io.sleep(...) path does suspend normally for ~200ms. I also tested it with an explicitly initialized std.Io.Threaded instance, and that also suspended normally.

I also tested the std.posix.system.nanosleep version, and that works correctly in Linux no-libc mode.

However, the std.c.nanosleep version does fail on Linux no-libc with:

error: dependency on libc must be explicitly specified in the build command

So I can confirm that std.posix.system.nanosleep is more portable than std.c.nanosleep for Linux no-libc builds, but I could not confirm that the old std.Io.sleep(...) implementation fails to suspend the OS thread on Zig 0.16.0.

Could you share the exact OS, architecture, Zig version, libc/linking mode, and runtime context where the old std.Io.sleep(...) behavior reproduces?

@vernonstinebaker

Copy link
Copy Markdown
Contributor Author

Thanks. Here is the exact environment where I saw the original high-CPU behavior.

  • macOS host: Mac mini, macOS 26.4, Apple Silicon M4, arm64
  • Zig: 0.16.0 installed via Homebrew
  • Linux devices: Radxa Zero 3W (aarch64) and OrangePi RV2 (riscv64)
  • Build/deploy shape: the Linux agents are cross-compiled on the Mac for bare-metal aarch64-linux-musl and riscv64-linux-musl targets, then run directly on the devices; no Docker/container layer involved
  • Runtime context: real long-running NullClaw agent/gateway processes, not a standalone microtest

The 100%+ CPU symptom manifested for me on both macOS and the Radxa. I do not recall for sure whether I saw the same issue on the OrangePi. In my testing at the time, integrating this nanosleep change normalized CPU on the affected hosts.

So the environments where I observed the problem were:

  • native macOS arm64
  • bare-metal aarch64 Linux built as aarch64-linux-musl

That is why I still believe the issue was real in production NullClaw runtimes even though it is not showing up in your Docker/Linux repro.

@vernonstinebaker

Copy link
Copy Markdown
Contributor Author

One additional note: the current PR head now uses std.posix.system.nanosleep(...) with EINTR retry, so I believe that addresses the Linux no-libc portability concern better than the older std.c.nanosleep(...) form did.

@DonPrus

DonPrus commented May 28, 2026

Copy link
Copy Markdown
Contributor

Thanks for the detailed environment notes.

One thing still looks a bit strange to me: I checked Zig 0.16.0’s implementation in lib/zig/std/Io/Threaded.zig, and under the hood std.Io.sleep(...) already uses clock_nanosleep(...) / nanosleep(...) on POSIX targets. So, at least from the Zig stdlib code path, std.Io.sleep should already be suspending the OS thread on macOS/Linux rather than spinning.

I also tried to reproduce the old behavior in Docker from Apple Silicon hosts, including MacBook M3/M4 setups, and I still cannot reproduce the high-CPU behavior there. In those environments the old std.Io.sleep(...) path appears to sleep normally.

That said, the current PR is small and the new implementation is explicit: on Linux/POSIX it calls std.posix.system.nanosleep(...) directly and retries on EINTR. This also avoids the previous std.c.nanosleep(...) portability issue for Linux no-libc-style builds. The change is isolated to compat.thread.sleep, keeps Windows/WASI on the existing path, and the test suite passes.

So my current read is:

  • I do not think we have proven that std.Io.sleep(...) itself was the root cause :)
  • The original production symptom may have depended on the real long-running agent/gateway runtime, cross-compiled target, host/device combination, or another polling loop around sleep.
  • But this PR is a reasonable defensive hardening change and gives us a simpler, direct POSIX sleep path.
  • I am okay merging it and validating on real devices/runtime, especially since the Docker repro does not seem to capture the environment where you observed the issue.

If the merged build keeps idle CPU normalized in the real agent/gateway process, that is good enough operationally, even if the exact underlying Zig/std.Io behavior remains unclear for me ))

@DonPrus
DonPrus merged commit 4eb8efc into nullclaw:main May 28, 2026
3 checks passed
@vernonstinebaker

Copy link
Copy Markdown
Contributor Author

I didn't dig into the Zig code. It is strange that it uses the same posix method under the hood but there seemed to be divergent behavior. That's curious indeed. There was a bit of churn as we migrated to Zig 0.16.0 so it's possible something else was causing the issue. I'm only sure that after I introduced that change in my local environments the CPU utilization became normal again.

@vernonstinebaker
vernonstinebaker deleted the fix/compat-thread-sleep branch May 30, 2026 02:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants