What happened?
Any LSP tool call whose server response contains non-ASCII characters (e.g. CJK) silently returns an empty result, even though the language server produced a valid, complete response.
Concrete case: documentSymbol on a Markdown file whose headings are Chinese returns No document symbols found for <file>., while the identical call on an ASCII-only file returns the symbols. The same happens for any other LSP method whose payload carries non-ASCII text.
The language server is not at fault. Verified with strace that the server writes a correct, fully-framed response and the qwen client reads all of it; the client still treats the request as empty and reports [], with no error surfaced.
What did you expect to happen?
Non-ASCII LSP responses should be parsed normally. documentSymbol should return the symbols regardless of the script used in the document.
Root cause
packages/core/src/lsp/LspConnectionFactory.ts -> JsonRpcConnection.handleData() mixes two different length units: the Content-Length header is a byte count, but this.buffer is a JS string, so this.buffer.length counts UTF-16 code units.
Current source (verified against main):
// line 32
private buffer = '';
// line 163
this.buffer += chunk.toString('utf8'); // buffer becomes a string -> .length = UTF-16 code units
// lines 176-182
const contentLength = Number(lengthMatch[1]); // Content-Length = BYTES
const messageStart = headerEnd + 4;
const messageEnd = messageStart + contentLength;
if (this.buffer.length < messageEnd) { // code units compared against bytes
break; // never satisfied for non-ASCII -> message never parsed
}
Note the asymmetry inside the same file: the write path already uses bytes correctly.
// line 237
const header = `Content-Length: ${Buffer.byteLength(json, 'utf8')}\r\n\r\n`;
For any non-ASCII body, Buffer.byteLength(body, 'utf8') > body.length, so the reader's guard can never be satisfied. The pending request then hits DEFAULT_LSP_REQUEST_TIMEOUT_MS (15 s) and resolves to an empty result — a silent false negative (no exception is raised, so the timeout is indistinguishable from "no symbols").
Reproduction (unit-level, self-contained)
The snippet below reproduces the exact handleData() logic and needs no editor/LSP server:
// handleData() copied verbatim from packages/core/src/lsp/LspConnectionFactory.ts
function handleData(state, chunk) {
state.buffer += chunk.toString('utf8'); // 1. bytes decoded into a JS string
while (true) {
const headerEnd = state.buffer.indexOf('\r\n\r\n');
if (headerEnd === -1) break;
const m = /Content-Length:\s*(\d+)/i.exec(state.buffer.slice(0, headerEnd));
if (!m) { state.buffer = state.buffer.slice(headerEnd + 4); continue; }
const contentLength = Number(m[1]); // 2. bytes
const messageStart = headerEnd + 4;
const messageEnd = messageStart + contentLength;
if (state.buffer.length < messageEnd) return false; // 3. UTF-16 code units
state.buffer = state.buffer.slice(messageEnd);
return true;
}
return false;
}
const body = JSON.stringify({ jsonrpc: '2.0', id: 1, result: [{ name: 'H1: 标题甲', kind: 15 }] });
const bytes = Buffer.byteLength(body, 'utf8');
const frame = Buffer.from(`Content-Length: ${bytes}\r\n\r\n${body}`, 'utf8');
console.log('body bytes (Content-Length) =', bytes);
console.log('body .length (code units) =', body.length);
console.log('CJK ->', handleData({ buffer: '' }, frame) ? 'OK' : 'DROPPED -> pending hangs -> timeout -> []');
const ascii = JSON.stringify({ jsonrpc: '2.0', id: 2, result: [{ name: 'H1: Title', kind: 15 }] });
const ab = Buffer.byteLength(ascii, 'utf8');
console.log('ASCII->', handleData({ buffer: '' }, Buffer.from(`Content-Length: ${ab}\r\n\r\n${ascii}`)) ? 'OK' : 'DROPPED');
Output:
body bytes (Content-Length) = 70
body .length (code units) = 64
CJK -> DROPPED -> pending hangs -> timeout -> []
ASCII-> OK
Reproduction (end-to-end)
- Configure a language server that answers
textDocument/documentSymbol by reading the file from disk, without requiring a prior didOpen (marksman does this).
- Call
documentSymbol on an ASCII-only .md file -> symbols returned.
- Call
documentSymbol on an otherwise identical .md file whose headings are CJK -> empty result.
Observed with strace on both processes:
- server:
write(27, "Content-Length: 11815\r\n\r\n", 25) = 25, then a 11815-byte body
- client:
read(31, "{\"jsonrpc\":\"2.0\",\"id\":18,\"result\":[{\"name\":\"H1: Changelog\"...", 65536) = 11815 (every byte read)
- yet the client still resolves the request to
[]
ASCII vs CJK A/B on otherwise identical files:
# Alpha Section / ## Beta Section -> 3 symbols
# 标题甲 / ## 标题乙 -> empty
As a control, documentSymbol via rust-analyzer on an ASCII .rs file returns symbols normally (so routing/registration is fine — only non-ASCII framing is broken).
Suggested fix
Accumulate the frame in a Buffer and compare byte lengths:
private buffer: Buffer = Buffer.alloc(0);
// inside handleData():
this.buffer = Buffer.concat([this.buffer, chunk]);
const headerEnd = this.buffer.indexOf('\r\n\r\n');
// ...
if (this.buffer.byteLength < messageEnd) break;
const body = this.buffer.subarray(messageStart, messageEnd).toString('utf8');
this.buffer = this.buffer.subarray(messageEnd);
Alternatively, keep the string accumulator but compare against Buffer.byteLength(bodyReadSoFar, 'utf8').
Client information
- qwen code version: 0.24.0
- Platform: Linux (x86_64)
- Launch:
qwen serve ... --experimental-lsp (ACP child)
- The same logic ships in the published bundle (
chunk-CNG7ILRW.js)
Anything else we need to know?
Separately observed in the same session: two complete LSP server sets are started from a single session (6 servers each, e.g. two marksman, two clangd, ...), i.e. one ACP child ends up with two NativeLspService instances. Both remain strongly referenced by the parent for the process lifetime. A likely cause is createWorkspaceMcpDiscoveryConfig() calling loadCliConfig() with ...this.argv (which keeps --experimental-lsp) without forcing experimentalLsp: false, unlike the per-session config path. Happy to file that as a separate issue if useful.
What happened?
Any LSP tool call whose server response contains non-ASCII characters (e.g. CJK) silently returns an empty result, even though the language server produced a valid, complete response.
Concrete case:
documentSymbolon a Markdown file whose headings are Chinese returnsNo document symbols found for <file>., while the identical call on an ASCII-only file returns the symbols. The same happens for any other LSP method whose payload carries non-ASCII text.The language server is not at fault. Verified with
stracethat the server writes a correct, fully-framed response and the qwen client reads all of it; the client still treats the request as empty and reports[], with no error surfaced.What did you expect to happen?
Non-ASCII LSP responses should be parsed normally.
documentSymbolshould return the symbols regardless of the script used in the document.Root cause
packages/core/src/lsp/LspConnectionFactory.ts->JsonRpcConnection.handleData()mixes two different length units: theContent-Lengthheader is a byte count, butthis.bufferis a JS string, sothis.buffer.lengthcounts UTF-16 code units.Current source (verified against
main):Note the asymmetry inside the same file: the write path already uses bytes correctly.
For any non-ASCII body,
Buffer.byteLength(body, 'utf8') > body.length, so the reader's guard can never be satisfied. The pending request then hitsDEFAULT_LSP_REQUEST_TIMEOUT_MS(15 s) and resolves to an empty result — a silent false negative (no exception is raised, so the timeout is indistinguishable from "no symbols").Reproduction (unit-level, self-contained)
The snippet below reproduces the exact
handleData()logic and needs no editor/LSP server:Output:
Reproduction (end-to-end)
textDocument/documentSymbolby reading the file from disk, without requiring a priordidOpen(marksman does this).documentSymbolon an ASCII-only.mdfile -> symbols returned.documentSymbolon an otherwise identical.mdfile whose headings are CJK -> empty result.Observed with
straceon both processes:write(27, "Content-Length: 11815\r\n\r\n", 25) = 25, then a 11815-byte bodyread(31, "{\"jsonrpc\":\"2.0\",\"id\":18,\"result\":[{\"name\":\"H1: Changelog\"...", 65536) = 11815(every byte read)[]ASCII vs CJK A/B on otherwise identical files:
# Alpha Section/## Beta Section-> 3 symbols# 标题甲/## 标题乙-> emptyAs a control,
documentSymbolvia rust-analyzer on an ASCII.rsfile returns symbols normally (so routing/registration is fine — only non-ASCII framing is broken).Suggested fix
Accumulate the frame in a
Bufferand compare byte lengths:Alternatively, keep the string accumulator but compare against
Buffer.byteLength(bodyReadSoFar, 'utf8').Client information
qwen serve ... --experimental-lsp(ACP child)chunk-CNG7ILRW.js)Anything else we need to know?
Separately observed in the same session: two complete LSP server sets are started from a single session (6 servers each, e.g. two
marksman, twoclangd, ...), i.e. one ACP child ends up with twoNativeLspServiceinstances. Both remain strongly referenced by the parent for the process lifetime. A likely cause iscreateWorkspaceMcpDiscoveryConfig()callingloadCliConfig()with...this.argv(which keeps--experimental-lsp) without forcingexperimentalLsp: false, unlike the per-session config path. Happy to file that as a separate issue if useful.