Skip to content

bug: CLI agent crashes on Chinese input containing spaces / CLI Agent 在中文含空格输入时崩溃 #2984

Description

@silvermissile

Summary

The CLI interactive agent (zeroclaw agent) crashes with Error reading input: stream did not contain valid UTF-8 when the user types Chinese (CJK) text that contains ASCII spaces. Input without spaces works fine.

概要

CLI 交互模式(zeroclaw agent)在用户输入包含空格的中文文本时崩溃,报错 Error reading input: stream did not contain valid UTF-8。不含空格的中文输入正常工作。


Reproduction / 复现步骤

# Enter the CLI agent (e.g. via kubectl exec or direct terminal)
zeroclaw agent

# Input WITHOUT spaces — works fine ✅
> 今天调度系统有多少失败的任务?

# Input WITH spaces — crashes ❌
> 今天调度系统有 多少失败的 任务?
Error reading input: stream did not contain valid UTF-8
# (agent exits immediately)

Root Cause / 根因分析

The interactive loop in src/agent/loop_.rs uses std::io::stdin().read_line(), which requires the entire input to be valid UTF-8. If any byte sequence is not valid UTF-8, it returns an error and the loop breaks (exits).

交互循环使用了 std::io::stdin().read_line(),该方法要求输入必须是合法 UTF-8。任何无效字节序列都会导致错误并退出循环。

When running through a PTY chain (e.g. kubectl exec -it, SSH, or remote terminals), the transport layer may split data frames at space (0x20) boundaries. If a multi-byte UTF-8 character (Chinese characters are 3 bytes each in UTF-8) is being transmitted and the frame boundary happens to interrupt the byte sequence, Rust's BufRead implementation detects incomplete UTF-8 and returns an InvalidData error immediately, rather than waiting for the remaining bytes.

当通过 PTY 链路运行时(如 kubectl exec -it、SSH 或远程终端),传输层可能在空格(0x20)边界处拆分数据帧。如果多字节 UTF-8 字符(中文字符在 UTF-8 中为 3 字节)正在传输时被帧边界切断,Rust 的 BufRead 实现会检测到不完整的 UTF-8 并立即返回 InvalidData 错误,而不是等待后续字节到达。

This explains why:

  • No spaces → Chinese bytes are contiguous, no frame split mid-character → works
  • With spaces → PTY may split at space boundary, possibly mid-character → crash

这解释了为什么:

  • 无空格 → 中文字节连续,不会在字符中间拆帧 → 正常
  • 有空格 → PTY 可能在空格边界拆帧,可能切断字符 → 崩溃

Affected Code / 受影响代码

// src/agent/loop_.rs (line ~3057)
let mut input = String::new();
match std::io::stdin().read_line(&mut input) {
    Ok(0) => break,
    Ok(_) => {}
    Err(e) => {
        eprintln!("\nError reading input: {e}\n");
        break; // ← exits the entire agent loop
    }
}

Proposed Fix / 修复方案

Replace stdin().read_line() with byte-level stdin().lock().read_until(b'\n') followed by String::from_utf8_lossy(). This reads raw bytes until newline without UTF-8 validation, then does lossy conversion (replacing any truly invalid bytes with U+FFFD instead of crashing).

将 stdin().read_line() 替换为字节级的 stdin().lock().read_until(b'\n'),然后使用 String::from_utf8_lossy() 转换。这会读取原始字节直到换行符,不进行 UTF-8 校验,然后做有损转换(将真正无效的字节替换为 U+FFFD 而非崩溃)。

Additionally, set ENV LANG=C.UTF-8 in the Dockerfile as defense-in-depth.

同时在 Dockerfile 中设置 ENV LANG=C.UTF-8 作为纵深防御。

Environment / 环境

  • ZeroClaw v0.1.7+ (current master)
  • Running in Docker container (Debian trixie-slim, user nobody)
  • Accessed via kubectl exec -it (PTY over WebSocket)
  • Also reproducible via SSH to containers
  • Any CJK input with embedded ASCII spaces triggers the issue

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions