Description
Problem
Through extensive testing on Kimi K2.5 (SGLang), we identified three critical issues with the current Shell tool that significantly degrade the Agent's performance on Windows, particularly during the initial pass (pass-1) of command generation:
1. Ambiguous Shell Identity in Prompt
The original tool context failed to establish a clear PowerShell identity to the model. Even when the Agent explicitly knew it was running on Windows (via os_kind: "Windows"), the model would still default to bash-style commands (ls -la, cat file, && chains) because the prompt did not strongly anchor the PowerShell context. This led to unnecessary correction cycles.
Example failure pattern:
User request: List files in current directory
Agent thought: "I'm on Windows, I should use Shell tool..."
Generated command: `ls -la` ← bash-style, fails in PowerShell
2. Misleading Examples (CMD/PS Hybrid Confusion)
The original powershell.md contained ambiguous guidance and mixed CMD-style examples (dir, set, copy, findstr) with a few PowerShell cmdlets. This created a hybrid confusion: the model would assume PowerShell supports CMD syntax directly, leading to commands like:
# Commands that fail or behave unexpectedly in PowerShell
dir /b ← CMD-style switch fails
findstr "pattern" file ← Works but inconsistent with PS idioms
copy file1 file2 ← Works but misses PS-native alternatives
These silent incompatibilities caused the error rate to spike, especially when the model tried to chain commands using syntax that worked in examples but failed in actual execution.
3. Version-Blind Context (The && / || Trap)
PowerShell 5.1 (built into Windows) and 7+ (modern cross-platform) have breaking syntax differences:
| Operator |
PS 5.1 |
PS 7+ |
&& |
❌ Not supported |
✅ Supported |
|| |
❌ Not supported |
✅ Supported |
cmd1; cmd2 |
✅ Sequential |
✅ Sequential |
cmd1; if ($?) { cmd2 } |
✅ Conditional |
✅ Conditional |
The original tool context provided no version information, leaving the model blind to these critical differences. When generating commands like python test.py && echo "Success", the model had a ~50% chance of producing invalid syntax on PS 5.1 systems, causing immediate failures.
Impact: Without shell_version in the environment, the pass-1 command error rate was significantly higher, requiring multi-turn corrections.
Solution
This PR addresses all three issues through targeted prompt engineering and enhanced environment awareness:
1. Clear PowerShell Identity Establishment
New context files (powershell5.md, powershell7.md) open with an explicit identity declaration:
**⚠️ CRITICAL: You are running Windows PowerShell 5.1**
**✅ You are running PowerShell 7+**
This strong contextual anchor immediately orients the model to the correct command paradigm, eliminating the bash-default tendency.
2. Unambiguous Syntax Guidance
- Removed all CMD-style examples and ambiguous aliases
- Added explicit compatibility tables showing forbidden vs. required syntax
- Provided native PowerShell idioms (
Get-ChildItem, Select-String, Test-Path) with proper parameter usage
- Clear examples of correct chaining patterns for each version
3. Version-Aware Context Selection
Enhanced Environment Detection (environment.py):
- Added
shell_version field via runtime detection ($PSVersionTable.PSVersion.ToString())
- Prioritizes
pwsh.exe (PS7+) over legacy powershell.exe (PS5.1)
- Dynamic context selection via
_get_shell_key() method
Impact on Agent behavior:
| Before |
After |
Model sees: os_kind: "Windows" |
Model sees: os_kind: "Windows", shell_version: "5.1" |
| Uses ambiguous generic prompt |
Receives targeted PS 5.1 compatibility context |
&& chains fail silently |
Substitutes with ; if ($?) { ... } pattern |
Screenshots
Environment Detection
| BEFORE |
AFTER |
Only os_kind, shell_name |
Added shell_version with runtime version detection |
Tool Context Rendering
| BEFORE |
AFTER |
 |
 |
Generic powershell.md with mixed CMD/PS examples |
Version-specific context with clear identity and syntax rules |
Command Generation Quality
| BEFORE |
AFTER |
 |
 |
bash-style defaults, CMD syntax, &&/|| failures |
Native PowerShell idioms, version-compatible chaining |
Backward Compatibility
- Unix/Linux: No changes - continues using
bash.md
- Windows PS 5.1: Receives improved, more accurate context
- Windows PS 7+: Receives modern context enabling full operator support
Summary: By adding zero-overhead environment awareness and targeted prompt engineering, this PR significantly improves the Agent's shell command reliability on Windows, reducing the need for multi-turn corrections and improving the overall user experience.
Additional information
No response
Description
Problem
Through extensive testing on Kimi K2.5 (SGLang), we identified three critical issues with the current Shell tool that significantly degrade the Agent's performance on Windows, particularly during the initial pass (pass-1) of command generation:
1. Ambiguous Shell Identity in Prompt
The original tool context failed to establish a clear PowerShell identity to the model. Even when the Agent explicitly knew it was running on Windows (via
os_kind: "Windows"), the model would still default to bash-style commands (ls -la,cat file,&&chains) because the prompt did not strongly anchor the PowerShell context. This led to unnecessary correction cycles.Example failure pattern:
2. Misleading Examples (CMD/PS Hybrid Confusion)
The original
powershell.mdcontained ambiguous guidance and mixed CMD-style examples (dir,set,copy,findstr) with a few PowerShell cmdlets. This created a hybrid confusion: the model would assume PowerShell supports CMD syntax directly, leading to commands like:These silent incompatibilities caused the error rate to spike, especially when the model tried to chain commands using syntax that worked in examples but failed in actual execution.
3. Version-Blind Context (The
&&/||Trap)PowerShell 5.1 (built into Windows) and 7+ (modern cross-platform) have breaking syntax differences:
&&||cmd1; cmd2cmd1; if ($?) { cmd2 }The original tool context provided no version information, leaving the model blind to these critical differences. When generating commands like
python test.py && echo "Success", the model had a ~50% chance of producing invalid syntax on PS 5.1 systems, causing immediate failures.Impact: Without
shell_versionin the environment, the pass-1 command error rate was significantly higher, requiring multi-turn corrections.Solution
This PR addresses all three issues through targeted prompt engineering and enhanced environment awareness:
1. Clear PowerShell Identity Establishment
New context files (
powershell5.md,powershell7.md) open with an explicit identity declaration:This strong contextual anchor immediately orients the model to the correct command paradigm, eliminating the bash-default tendency.
2. Unambiguous Syntax Guidance
Get-ChildItem,Select-String,Test-Path) with proper parameter usage3. Version-Aware Context Selection
Enhanced Environment Detection (
environment.py):shell_versionfield via runtime detection ($PSVersionTable.PSVersion.ToString())pwsh.exe(PS7+) over legacypowershell.exe(PS5.1)_get_shell_key()methodImpact on Agent behavior:
os_kind: "Windows"os_kind: "Windows",shell_version: "5.1"&&chains fail silently; if ($?) { ... }patternScreenshots
Environment Detection
os_kind,shell_nameshell_versionwith runtime version detectionTool Context Rendering
powershell.mdwith mixed CMD/PS examplesCommand Generation Quality
&&/||failuresBackward Compatibility
bash.mdSummary: By adding zero-overhead environment awareness and targeted prompt engineering, this PR significantly improves the Agent's shell command reliability on Windows, reducing the need for multi-turn corrections and improving the overall user experience.
Additional information
No response