Description
A cron job of job_type: "agent" that never exits will block every other cron job for as long as it hangs, because cron jobs are dispatched serially on a single scheduler thread and the default agent timeout is 0 (no timeout).
Three separate defects combine to produce this:
1. scheduler.agent_timeout_secs defaults to 0 = no timeout
src/config_types.zig:343-349
pub const SchedulerConfig = struct {
enabled: bool = true,
max_tasks: u32 = 64,
max_concurrent: u32 = 4,
/// Hard timeout for cron agent subprocess execution. 0 = no timeout.
agent_timeout_secs: u64 = 0,
};
The kill machinery exists and works, but at the default it is unreachable: hasTimeoutExpired short-circuits on 0 (src/agent_runner.zig:32), and the drain loop then polls with poll_ms = -1 — block indefinitely — until both pipes hit EOF (src/agent_runner.zig:86-89), which is its only other exit condition.
This is an asymmetry within one config file: shell-type cron jobs are hard-bounded at 60s (DEFAULT_CRON_SHELL_TIMEOUT_NS, src/cron.zig:19), so agent-type jobs are the only unbounded kind.
The knob was also undocumented — it appeared in no docs/ page, the README, or config.example.json, so an operator had no way to know it existed. (Documented in #1032.)
2. scheduler.max_concurrent is parsed but never enforced
max_concurrent (src/config_types.zig:346) is parsed (src/config_parse.zig:1716) and printed by nullclaw cron status (src/status.zig:242), but src/cron.zig never reads it. CronJob has no running / in_flight field and tick has no skip-if-busy or queue logic, so the cron scheduler is unconditionally serial. A config field that silently does nothing is its own trap.
3. The timeout, when set, kills only the direct child
terminateChildHard (src/agent_runner.zig:140-151) sends SIGKILL to the direct child only, and the agent child is spawned without child.pgid (src/agent_runner.zig:255-259). The MCP stdio child does set it, with a comment naming this exact failure mode (src/mcp.zig:112):
// Put launchers and their descendants in one group so a timed-out
// MCP request cannot strand wrappers such as flock, ssh, or npx.
child.pgid = 0;
So a grandchild that inherits the job's output pipe survives the kill. Because poll_ms becomes -1 once timed_out is true (src/agent_runner.zig:86), the drain loop can then block past the timeout. agent_timeout_secs is therefore a mitigation, not a guarantee.
Expected behavior
- Cron agent jobs are bounded by default; a job that hangs cannot starve the rest of the schedule.
max_concurrent either does something, or is removed. A silently-ignored setting should not be surfaced by nullclaw cron status.
- A timeout terminates the job's whole process group, so descendants cannot hold the scheduler past the deadline.
Steps to reproduce
Conceptually, on a host with any cron agent job:
- Configure a cron
agent job (e.g. one that calls an MCP tool).
- Leave
scheduler.agent_timeout_secs unset (default 0).
- Make the tool call hang — an MCP server that accepts the connection and never responds is sufficient.
- Observe the scheduler: subsequent cron jobs stop being evaluated entirely, because
tick (src/cron.zig:897) awaits the in-flight job before examining the next one, holding the shared-scheduler mutex for the duration.
kill the hung child. tick returns, the drain loop's EOF condition is met, and the schedule resumes.
Observed in production: a fleet heartbeat-monitor job hung for 2h03m, during which the heartbeat job on the same scheduler did not fire. Killing the child restored the schedule immediately.
Note the schedule-wide impact is the real hazard, not the single stuck job — a monitoring job whose purpose is to detect missing heartbeats became itself the thing that stopped them.
Version
7757dd21 (main)
OS
Linux (reproduced on aarch64 bare-metal; the mechanism is platform-independent)
Description
A cron job of
job_type: "agent"that never exits will block every other cron job for as long as it hangs, because cron jobs are dispatched serially on a single scheduler thread and the default agent timeout is0(no timeout).Three separate defects combine to produce this:
1.
scheduler.agent_timeout_secsdefaults to0= no timeoutsrc/config_types.zig:343-349The kill machinery exists and works, but at the default it is unreachable:
hasTimeoutExpiredshort-circuits on0(src/agent_runner.zig:32), and the drain loop then polls withpoll_ms = -1— block indefinitely — until both pipes hit EOF (src/agent_runner.zig:86-89), which is its only other exit condition.This is an asymmetry within one config file: shell-type cron jobs are hard-bounded at 60s (
DEFAULT_CRON_SHELL_TIMEOUT_NS,src/cron.zig:19), so agent-type jobs are the only unbounded kind.The knob was also undocumented — it appeared in no
docs/page, the README, orconfig.example.json, so an operator had no way to know it existed. (Documented in #1032.)2.
scheduler.max_concurrentis parsed but never enforcedmax_concurrent(src/config_types.zig:346) is parsed (src/config_parse.zig:1716) and printed bynullclaw cron status(src/status.zig:242), butsrc/cron.zignever reads it.CronJobhas norunning/in_flightfield andtickhas no skip-if-busy or queue logic, so the cron scheduler is unconditionally serial. A config field that silently does nothing is its own trap.3. The timeout, when set, kills only the direct child
terminateChildHard(src/agent_runner.zig:140-151) sendsSIGKILLto the direct child only, and the agent child is spawned withoutchild.pgid(src/agent_runner.zig:255-259). The MCP stdio child does set it, with a comment naming this exact failure mode (src/mcp.zig:112):So a grandchild that inherits the job's output pipe survives the kill. Because
poll_msbecomes-1oncetimed_outis true (src/agent_runner.zig:86), the drain loop can then block past the timeout.agent_timeout_secsis therefore a mitigation, not a guarantee.Expected behavior
max_concurrenteither does something, or is removed. A silently-ignored setting should not be surfaced bynullclaw cron status.Steps to reproduce
Conceptually, on a host with any cron agent job:
agentjob (e.g. one that calls an MCP tool).scheduler.agent_timeout_secsunset (default0).tick(src/cron.zig:897) awaits the in-flight job before examining the next one, holding the shared-scheduler mutex for the duration.killthe hung child.tickreturns, the drain loop's EOF condition is met, and the schedule resumes.Observed in production: a fleet heartbeat-monitor job hung for 2h03m, during which the heartbeat job on the same scheduler did not fire. Killing the child restored the schedule immediately.
Note the schedule-wide impact is the real hazard, not the single stuck job — a monitoring job whose purpose is to detect missing heartbeats became itself the thing that stopped them.
Version
7757dd21(main)OS
Linux (reproduced on aarch64 bare-metal; the mechanism is platform-independent)