Summary
The CLI interactive agent (zeroclaw agent) crashes with Error reading input: stream did not contain valid UTF-8 when the user types Chinese (CJK) text that contains ASCII spaces. Input without spaces works fine.
概要
CLI 交互模式(zeroclaw agent)在用户输入包含空格的中文文本时崩溃,报错 Error reading input: stream did not contain valid UTF-8。不含空格的中文输入正常工作。
Reproduction / 复现步骤
# Enter the CLI agent (e.g. via kubectl exec or direct terminal)
zeroclaw agent
# Input WITHOUT spaces — works fine ✅
> 今天调度系统有多少失败的任务?
# Input WITH spaces — crashes ❌
> 今天调度系统有 多少失败的 任务?
Error reading input: stream did not contain valid UTF-8
# (agent exits immediately)
Root Cause / 根因分析
The interactive loop in src/agent/loop_.rs uses std::io::stdin().read_line(), which requires the entire input to be valid UTF-8. If any byte sequence is not valid UTF-8, it returns an error and the loop breaks (exits).
交互循环使用了 std::io::stdin().read_line(),该方法要求输入必须是合法 UTF-8。任何无效字节序列都会导致错误并退出循环。
When running through a PTY chain (e.g. kubectl exec -it, SSH, or remote terminals), the transport layer may split data frames at space (0x20) boundaries. If a multi-byte UTF-8 character (Chinese characters are 3 bytes each in UTF-8) is being transmitted and the frame boundary happens to interrupt the byte sequence, Rust's BufRead implementation detects incomplete UTF-8 and returns an InvalidData error immediately, rather than waiting for the remaining bytes.
当通过 PTY 链路运行时(如 kubectl exec -it、SSH 或远程终端),传输层可能在空格(0x20)边界处拆分数据帧。如果多字节 UTF-8 字符(中文字符在 UTF-8 中为 3 字节)正在传输时被帧边界切断,Rust 的 BufRead 实现会检测到不完整的 UTF-8 并立即返回 InvalidData 错误,而不是等待后续字节到达。
This explains why:
- No spaces → Chinese bytes are contiguous, no frame split mid-character → works
- With spaces → PTY may split at space boundary, possibly mid-character → crash
这解释了为什么:
- 无空格 → 中文字节连续,不会在字符中间拆帧 → 正常
- 有空格 → PTY 可能在空格边界拆帧,可能切断字符 → 崩溃
Affected Code / 受影响代码
// src/agent/loop_.rs (line ~3057)
let mut input = String::new();
match std::io::stdin().read_line(&mut input) {
Ok(0) => break,
Ok(_) => {}
Err(e) => {
eprintln!("\nError reading input: {e}\n");
break; // ← exits the entire agent loop
}
}
Proposed Fix / 修复方案
Replace stdin().read_line() with byte-level stdin().lock().read_until(b'\n') followed by String::from_utf8_lossy(). This reads raw bytes until newline without UTF-8 validation, then does lossy conversion (replacing any truly invalid bytes with U+FFFD instead of crashing).
将 stdin().read_line() 替换为字节级的 stdin().lock().read_until(b'\n'),然后使用 String::from_utf8_lossy() 转换。这会读取原始字节直到换行符,不进行 UTF-8 校验,然后做有损转换(将真正无效的字节替换为 U+FFFD 而非崩溃)。
Additionally, set ENV LANG=C.UTF-8 in the Dockerfile as defense-in-depth.
同时在 Dockerfile 中设置 ENV LANG=C.UTF-8 作为纵深防御。
Environment / 环境
- ZeroClaw v0.1.7+ (current master)
- Running in Docker container (Debian trixie-slim, user
nobody)
- Accessed via
kubectl exec -it (PTY over WebSocket)
- Also reproducible via SSH to containers
- Any CJK input with embedded ASCII spaces triggers the issue
Summary
The CLI interactive agent (
zeroclaw agent) crashes withError reading input: stream did not contain valid UTF-8when the user types Chinese (CJK) text that contains ASCII spaces. Input without spaces works fine.概要
CLI 交互模式(
zeroclaw agent)在用户输入包含空格的中文文本时崩溃,报错Error reading input: stream did not contain valid UTF-8。不含空格的中文输入正常工作。Reproduction / 复现步骤
Root Cause / 根因分析
The interactive loop in
src/agent/loop_.rsusesstd::io::stdin().read_line(), which requires the entire input to be valid UTF-8. If any byte sequence is not valid UTF-8, it returns an error and the loop breaks (exits).交互循环使用了
std::io::stdin().read_line(),该方法要求输入必须是合法 UTF-8。任何无效字节序列都会导致错误并退出循环。When running through a PTY chain (e.g.
kubectl exec -it, SSH, or remote terminals), the transport layer may split data frames at space (0x20) boundaries. If a multi-byte UTF-8 character (Chinese characters are 3 bytes each in UTF-8) is being transmitted and the frame boundary happens to interrupt the byte sequence, Rust'sBufReadimplementation detects incomplete UTF-8 and returns anInvalidDataerror immediately, rather than waiting for the remaining bytes.当通过 PTY 链路运行时(如
kubectl exec -it、SSH 或远程终端),传输层可能在空格(0x20)边界处拆分数据帧。如果多字节 UTF-8 字符(中文字符在 UTF-8 中为 3 字节)正在传输时被帧边界切断,Rust 的BufRead实现会检测到不完整的 UTF-8 并立即返回InvalidData错误,而不是等待后续字节到达。This explains why:
这解释了为什么:
Affected Code / 受影响代码
Proposed Fix / 修复方案
Replace
stdin().read_line()with byte-levelstdin().lock().read_until(b'\n')followed byString::from_utf8_lossy(). This reads raw bytes until newline without UTF-8 validation, then does lossy conversion (replacing any truly invalid bytes with U+FFFD instead of crashing).将
stdin().read_line()替换为字节级的stdin().lock().read_until(b'\n'),然后使用String::from_utf8_lossy()转换。这会读取原始字节直到换行符,不进行 UTF-8 校验,然后做有损转换(将真正无效的字节替换为 U+FFFD 而非崩溃)。Additionally, set
ENV LANG=C.UTF-8in the Dockerfile as defense-in-depth.同时在 Dockerfile 中设置
ENV LANG=C.UTF-8作为纵深防御。Environment / 环境
nobody)kubectl exec -it(PTY over WebSocket)