Skip to content

Windows: shell wrapper via powershell.exe -Command mojibakes UTF-8 files on non-English locales (PowerShell 5.1 Get-Content defaults to ANSI codepage) #23044

Description

@cyesuta

What version of Codex CLI is running?

@openai/codex-sdk 0.130.0 (programmatic JSON-RPC via codex app-server; not the codex CLI binary itself). Same wrapper behavior is observable on the codex CLI build of the same vintage.

What subscription do you have?

ChatGPT Plus

Which model were you using?

gpt-5.5

What platform is your computer?

Windows 10/11, Chinese (Simplified) zh-CN system locale. PowerShell version: Windows PowerShell 5.1 (the default Windows ships with — C:\Windows\System32\WindowsPowerShell\v1.0\powershell.exe).

What terminal emulator and version are you using (if applicable)?

custom Tauri-based editor that embeds codex app-server via the TypeScript SDK. The same mojibake reproduces from a plain Windows Terminal session driving codex CLI directly.

Codex doctor report

not available (using @openai/codex-sdk programmatically; `codex doctor` is the CLI-only entrypoint)

What issue are you seeing?

On non-English Windows (any CJK / Arabic / Cyrillic locale), the codex shell tool produces mojibake whenever it reads a UTF-8-without-BOM text file. The model receives the corrupted text and silently answers as if the corrupted text were the file's actual content — the failure mode is "confident wrong answer," not an exception, so it can be mistaken for a model hallucination.

Concrete observation: a skill file containing the literal string 貝克街15號 (UTF-8 bytes E8 B2 9D E5 85 8B E8 A1 97 31 35 E8 99 9F) was read by the model as the mojibake 璨濆厠琛?5铏. The model's final user-facing answer became 貝克街 5號 — note that "1" is gone and a space appears, because the mojibake bytes still form valid (but wrong) Unicode codepoints which the model then tried to reason about.

Root cause

The codex Windows shell wrapper invokes:


powershell.exe -Command "<model_emitted_command>"

(seen verbatim in our shell-tool execution logs: "C:\\WINDOWS\\System32\\WindowsPowerShell\\v1.0\\powershell.exe" -Command "Get-Content -Path E:\\...")

The wrapped command commonly contains Get-Content, or aliases like cat (which is a PowerShell alias for Get-Content). On Windows PowerShell 5.1, Get-Content defaults to [System.Text.Encoding]::Default when -Encoding is not specified — and Encoding::Default is the system ANSI codepage, NOT UTF-8:

Windows locale ANSI codepage Behavior reading UTF-8 file
zh-CN (Simplified Chinese) CP936 (GBK) mojibake
zh-TW / zh-HK (Traditional Chinese) CP950 (Big5) mojibake
ja-JP (Japanese) CP932 (Shift-JIS) mojibake
ko-KR (Korean) CP949 mojibake
en-US (English) CP1252 hidden: CP1252 ≡ ASCII for 0x00–0x7F, so pure-ASCII files appear fine; only files containing Latin-1 high bits manifest. Most en-US testing never hits a manifesting file.

This is the well-known PowerShell 5.1 default-encoding behavior that Microsoft fixed in PowerShell 7+ (which defaults to UTF-8 across cmdlets). Codex on Windows ends up sitting on top of the broken 5.1 default because powershell.exe is the only PowerShell guaranteed to be installed on every Windows machine.

The existing prefix_powershell_script_with_utf8() utility is incomplete

codex-rs/shell-command/src/powershell.rs already has a helper that prepends:

[Console]::OutputEncoding=[System.Text.Encoding]::UTF8;

That fixes the stdout-write side, but it does NOT change how Get-Content reads files. To fix the read side, the prefix must additionally set:

$PSDefaultParameterValues['Get-Content:Encoding'] = 'utf8'
$PSDefaultParameterValues['Set-Content:Encoding'] = 'utf8'
$PSDefaultParameterValues['Add-Content:Encoding'] = 'utf8'
$PSDefaultParameterValues['Out-File:Encoding']    = 'utf8'

So even the codex code paths that DO call the existing UTF-8 prefix utility are still letting Get-Content mojibake on read.

What steps can reproduce the bug?

System: Windows 10 or 11, set to a non-Latin-1 locale (zh-CN reproduces; zh-TW / ja-JP / ko-KR equally affected; en-US likely will NOT reproduce due to CP1252/ASCII overlap).

  1. Create a UTF-8-without-BOM text file containing non-ASCII text:

    "貝克街15號" | Out-File -Encoding utf8NoBOM C:\tmp\repro.md

    Verify with Format-Hex -Path C:\tmp\repro.md — the bytes should be
    E8 B2 9D E5 85 8B E8 A1 97 31 35 E8 99 9F (clean UTF-8).

  2. In a codex thread (CLI or app-server, any thread ID), ask:

    What's inside C:\tmp\repro.md?
    
  3. The model will emit a read command (typically cat C:\tmp\repro.md or Get-Content C:\tmp\repro.md).

  4. Codex wraps it as:

    powershell.exe -Command "Get-Content -Path C:\tmp\repro.md"
    
  5. PowerShell 5.1 decodes the UTF-8 file bytes through CP936 (or the relevant ANSI codepage) and returns mojibake to the model.

  6. The model's final user-facing answer reports the mojibake content as if it were the file content — e.g. our run produced 貝克街 5號 (the digit 1 is gone, a space appears) when the actual file content is 貝克街15號. The corrupted bytes still represent valid Unicode codepoints, so the model never realizes the data is wrong.

This reproduces 100% of the time on a zh-CN Windows host when reading any UTF-8 file containing non-ASCII content. It does NOT reproduce on en-US Windows reading the same file (the bug is real but hidden by the CP1252/ASCII byte overlap).

What is the expected behavior?

codex should return the file's actual content (貝克街15號) to the model, not the ANSI-codepage-misinterpreted byte string. The end-user-facing answer should accordingly be based on the actual content.

The fix is to make codex's Windows shell wrapper produce UTF-8 reads regardless of the user's system ANSI codepage. Concretely, any of the following would resolve the issue (in order of robustness):

  1. Always inject full UTF-8 init in -Command prefix — Console + $PSDefaultParameterValues together — not optionally via the existing prefix utility but unconditionally on every wrapped command. ~6 statement-separated lines, all on one logical line via ;.
  2. Prefer pwsh.exe (PS 7+) over powershell.exe (PS 5.1) when present — PS 7+ defaults to UTF-8 for cmdlets and would dodge the issue entirely. Probe once at startup (where.exe pwsh), cache the result, fall back to powershell.exe only when absent.
  3. Switch the Windows wrapper to cmd.exe + chcp 65001 >nul && <cmd> — loses PowerShell's Unix-style aliases (cat, ls, which etc.), probably not worth the disruption.
  4. Prefer Git Bash / WSL bash when available — matches the Linux/macOS code path most directly; requires bash to be installed which is not universal on Windows.

Recommended combination: fix 1 unconditionally (cheap, handles legacy 5.1 thoroughly) plus fix 2 when pwsh is present (makes the problem disappear entirely for PS 7+ users).

Additional information

Workaround we currently use (not from inside codex)

Since the wrapper cannot be changed from outside codex, we write a user-level PowerShell profile at

%UserProfile%\Documents\WindowsPowerShell\Profile.ps1

containing the full UTF-8 init shown in "Expected behavior" above. PowerShell auto-loads this on any powershell.exe spawn that does NOT pass -NoProfile. The codex-rs/app-server shell wrapper (the one reproduced above) does not pass -NoProfile, so the profile applies and self-heals the bug on that code path.

This workaround does NOT work for the codex IDE Extension wrapper described in #17208, which DOES pass -NoProfile. That code path therefore still mojibakes on non-English Windows even with the profile in place — fixing it requires an in--Command UTF-8 init (fix 1 above); no external workaround is available.

Related issues

Why this has been hard to catch upstream

  • Locale-hidden: CP1252/ASCII overlap means pure-ASCII files on en-US Windows appear fine. en-US testing never manifests this bug.
  • No exception: The model "succeeds" with a confidently-wrong answer rather than failing loudly, so it can be mistaken for a model hallucination rather than a wire-level encoding bug.
  • Estimated impact: ~25% of Windows users globally (any CJK / Arabic / Cyrillic / Greek / Hebrew locale).

Source pointers for the fix

  • codex-rs/shell-command/src/powershell.rs — existing prefix_powershell_script_with_utf8() is the place to extend with $PSDefaultParameterValues defaults (or to add a probe for pwsh.exe).
  • codex-rs/app-server/ — the wrapper invocation site (where to ensure the extended prefix is always applied, not optional).

Happy to verify any candidate fix against our zh-CN test setup. Thanks for codex — this is the only sustained rough edge we've hit in Windows use.

Activity

added
CLIIssues related to the Codex CLI
windows-osIssues related to Codex on Windows systems
tool-callsIssues related to tool calling
app-serverIssues involving app server protocol or interfaces
on May 16, 2026
added
TUIIssues related to the terminal user interface: text input, menus and dialogs, and terminal display
and removed
CLIIssues related to the Codex CLI
app-serverIssues involving app server protocol or interfaces
TUIIssues related to the terminal user interface: text input, menus and dialogs, and terminal display
on May 16, 2026

gianlucafarias commented on May 17, 2026

@gianlucafarias

I have a tested fix ready in my fork for this PowerShell UTF-8 mojibake issue.

Branch:

What it changes:

  • Expands the PowerShell UTF-8 prefix to set:
    • Console Input/OutputEncoding
    • $OutputEncoding
    • $PSDefaultParameterValues for Get-Content/Set-Content/Add-Content/Out-File

Validation run locally:

  • cargo test -p codex-shell-command (136 passed)
  • cargo test -p codex-shell-command powershell::tests:: (6 passed)
  • cargo clippy -p codex-shell-command --tests -- -D warnings
  • cargo fmt --all -- --check

Could a maintainer invite me to open a PR for this fix?
I can also attach a patch file if preferred.

claell commented on Jun 7, 2026

@claell

Adding a few additional observations from a separate reproduction on Windows with GPT-5.5 in high effort mode. Filing on this issue rather than opening a new one, because the root cause appears to be the same as described in the body.

  1. The bug does not only produce wrong answers — it also causes Codex to write corrupted content back to the file. When Codex reads a file with the broken encoding, it sometimes treats the garbled output as a real defect and overwrites the file with its own interpretation, effectively writing ASCII or garbled content over the original. This is the variant that causes actual data loss.

  2. Frequency impression: on GPT-5.5 in high effort mode this has become a recurring issue across multiple sessions. I have the impression it became more frequent relatively recently — it was not happening at first — but I have not bisected a specific build, so I am presenting this as an observation rather than a regression claim.

  3. Cross-referencing False Turkish encoding corruption warnings during code review on Windows PowerShell #13755 for visibility — it tracks the same underlying read-encoding root cause from the false-positive-warning angle and is worth linking here.

dj-thank commented on Jul 12, 2026

@dj-thank

I can reproduce this on the current Codex Desktop build on Japanese Windows, and the attachment write path is not corrupting the file.

Environment

  • Codex Desktop 26.707.3748.0
  • bundled/session CLI 0.144.0-alpha.4
  • Windows 10.0.26200 x64
  • Windows PowerShell 5.1.26100.8655
  • system locale ja-JP, active ANSI code page 932

Evidence

Codex Desktop created a long-paste attachment as pasted-text.txt. I am not publishing the original because it contains private infrastructure details, but I verified the file independently:

  • 17,630 bytes
  • no BOM
  • strict UTF-8 decode succeeds
  • decoding and re-encoding as UTF-8 reproduces the original bytes exactly
  • no replacement characters or pre-existing mojibake markers

The rollout then invoked the equivalent of:

Get-Content -Raw -LiteralPath '<attachment>\pasted-text.txt'

with no -Encoding. Japanese text was mojibaked in the tool output. The controls are deterministic:

strict UTF-8 == Get-Content default        -> false
strict UTF-8 == Get-Content -Encoding UTF8 -> true

This isolates the failure to the Windows PowerShell 5.1 read boundary, not clipboard capture, attachment persistence, console rendering, or the model.

Current main at 9e552e9d15ba52bed7077d5357f3e18e330f8f38 still only sets [Console]::OutputEncoding; that cannot repair a string already decoded through CP932.

I prepared a current-main commit that preserves the existing best-effort console prefix and adds a separate best-effort default only for legacy PowerShell:

if ($PSVersionTable.PSVersion.Major -lt 6) {
    $PSDefaultParameterValues['Get-Content:Encoding'] = 'utf8'
}

The change intentionally does not set *:Encoding or write-cmdlet defaults, because Windows PowerShell 5.1 utf8 writes a BOM and would broaden behavior unnecessarily. An explicit user -Encoding still wins. The patch includes a Windows regression test that writes BOM-less Japanese UTF-8 and asserts an exact read through the prefixed powershell.exe command.

Behavioral verification passes on PS 5.1 / CP932, ConstrainedLanguage, and PowerShell 7.6.3. Rust formatting and diff checks pass. The full local Cargo test build is still pending because Windows Smart App Control blocks Cargo-generated unsigned build scripts on this machine (Code Integrity 3077), so CI validation would still be required.

This is the same fallback path discussed in #29085. Per the invitation-only contribution policy, I have not opened an unsolicited PR. If a maintainer would like this focused current-main patch, please invite the PR and I will push it.

Meir770ar commented on Aug 26, 2026

@Meir770ar

Another locale confirmation (Hebrew, ACP 1255), plus two measurements that I think
change the shape of the proposed fix. Filing here rather than opening a new issue since
the root cause is the same one described in the body.

TL;DR: the existing [Console]::OutputEncoding prefix does not actually work, and
PowerShell 7 alone does not fix this.
Both are load-bearing assumptions above.

1. The existing prefix is silently a no-op

The body says the current helper "fixes the stdout-write side". On my machine it does not,
because of how it is wrapped. This is the command Codex builds, quoted verbatim from its
own error output:

pwsh.exe -NoProfile -Command "try { [Console]::OutputEncoding=[System.Text.Encoding]::UTF8 } catch {}
Get-Content -Raw -LiteralPath .\sample.mjs"

The [Console]::OutputEncoding setter throws when no console is attached, which is the
case for a child spawned with piped stdio and windowsHide. The catch {} discards it and
the shell falls back to the inherited console codepage.

Verified by asking Codex to report its own state after that prefix has run:

Active code page: 862
[Console]::OutputEncoding.CodePage  ->  862
[Console]::InputEncoding.CodePage   ->  862
$PSVersionTable.PSVersion           ->  7.6.5

So the write side is unprotected too, not just the read side. Worth wrapping the fix in
something that surfaces failure rather than catch {} — this was invisible for months
precisely because that exception never reached anyone.

2. PowerShell 7 is not sufficient on its own

The body notes PS7 "defaults to UTF-8 across cmdlets", which implies moving to pwsh 7
resolves this. It does not. Controlled experiment, spawning each shell exactly the way
Codex does (Node spawn, piped stdio, windowsHide,
-NoLogo -NoProfile -NonInteractive -EncodedCommand), reading the same UTF-8 file:

shell console codepage result
Windows PowerShell 5.1 inherited (862) corrupt
Windows PowerShell 5.1 65001 corrupt — reproduces the exact bytes found in my rollout files
pwsh 7.6.5 inherited (862) corrupt (different signature: ✓ best-fits to √ via cp862)
pwsh 7.6.5 65001 clean

Row 2 is the useful one: it maps the lab back to the field. Row 3 is the one that matters
for the fix — pwsh 7 fixes the read default but still encodes its output to the
inherited console codepage, so $PSDefaultParameterValues alone will not close this. The
console codepage has to be set explicitly (SetConsoleOutputCP/SetConsoleCP at process
creation, or chcp 65001 as the first statement).

3. It also produces hard apply_patch failures, not only wrong answers

The body frames this as "confident wrong answer". On this machine it additionally breaks
editing outright, because the model builds patch context from the mangled text:

ERROR codex_core::tools::router: error=apply_patch verification failed:
Failed to find expected lines in <path>\<file>.mjs:
console.log('ג“ visual windows preserve overlapping narration context across cut boundaries');

The file contains ✓ there. ג + U+009C + “ is precisely the cp1255 decode of
E2 9C 93, stored in the rollout as c2 9c e2 80 9c.

Scale on one machine, across 349 rollout files: 155 apply_patch verification failed
occurrences, and 25,739 instances of a single signature (ג€, U+2014 through cp1255)
across 205 of 343 sessions.

Confirming what @claell reported in the second comment about write-back: Codex's own
~/.codex/config.toml accumulated corrupted Hebrew [projects.*] keys, so the corruption
reaches Codex's own config, not just user files.

4. New-ish failure mode: fabricated patch context when the shell is fully broken

While the shell was failing outright, Codex issued four consecutive speculative
apply_patch calls against a file it had never successfully read, inventing "pending",
"unverified", 'pending', "needs review" as the file's current contents. A failed read
degrading into invented patch context seems worth guarding independently of the encoding
fix.

Workaround, for anyone hitting this before the fix lands

Launch codex through a wrapper that sets the codepage first. It works because the console
is inherited down the whole chain (cmd -> codex.exe -> pwsh.exe):

@echo off
chcp 65001 >nul
"C:\path\to\codex.exe" %*

Verified end to end afterwards: shell read clean, apply_patch succeeded, and the bytes
written to disk were correct (e2 9c 93 and e2 80 94 preserved). Note this needs real
PowerShell 7 present as well — see below.

Caution for anyone installing PowerShell 7 to work around this

winget install --id Microsoft.PowerShell installs the MSIX/Store build (that manifest
is MSIX-only). Its app-execution alias in WindowsApps is a 0-byte reparse point, which
[windows] sandbox = "unelevated" cannot launch:

windows sandbox: CreateProcessAsUserW failed: 5 (Access is denied.)
| cmd=C:\Users\<user>\AppData\Local\Microsoft\WindowsApps\pwsh.exe -NoProfile -Command ...

Codex prefers that binary from PATH over Windows PowerShell 5.1, so installing PowerShell 7
the most obvious way takes the shell tool from "corrupts text" to "every command fails at
process creation". The MSI from GitHub releases installs a real executable and works. This
overlaps with #39276; might be worth skipping reparse-point binaries during shell discovery.

Environment: codex-cli 0.150.0-alpha.8 (also reproduced on 0.147.0-alpha.6.6 and
0.149.0-alpha.4.3), Windows 11 26200, he-IL, ACP 1255 / OEMCP 862.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingtool-callsIssues related to tool callingwindows-osIssues related to Codex on Windows systems

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions