Summary
Turn the navigate-code routing contract into regression cases so skill changes are driven by observed failures rather than by adding more prose.
The cases should verify which route is selected first, which follow-up evidence is required, and which claims are forbidden—for example, treating name-based text matches as semantic references.
Motivation
The skill is intentionally small. Without regression cases, future contributors have only manual judgment when deciding whether a new rule improves routing or bloats agent context.
This work protects the core product behavior: Semble for behavior-level discovery, Serena for symbols and relationships, and built-ins for literal search and focused reading.
Proposal
- Define routing cases for unknown behavior, exact text, known-file reasoning, references, implementations, renames, and architecture tracing.
- Score the first selected route, unnecessary calls, evidence quality, and prohibited claims separately from final-answer correctness.
- Include Serena project-activation behavior, 0-based Serena line handling, and Semble content-scope selection.
- Run cases against both Claude Code and Codex when automation permits; otherwise use a shared case format with client-specific runners.
- Require a failing case before materially expanding the skill body.
Acceptance criteria
Non-goals / out of scope
- Testing every Serena or Semble tool independently.
- Replacing repository-level correctness benchmarks.
- Requiring identical tool-call sequences from different models.
Edge cases
- A requested MCP tool is unavailable despite the plugin being installed.
- A symbol is exported by both application code and a dependency.
- Dynamic-language references cannot be proven by the language server.
- The correct route changes after semantic discovery identifies a concrete symbol.
Context
Summary
Turn the
navigate-coderouting contract into regression cases so skill changes are driven by observed failures rather than by adding more prose.The cases should verify which route is selected first, which follow-up evidence is required, and which claims are forbidden—for example, treating name-based text matches as semantic references.
Motivation
The skill is intentionally small. Without regression cases, future contributors have only manual judgment when deciding whether a new rule improves routing or bloats agent context.
This work protects the core product behavior: Semble for behavior-level discovery, Serena for symbols and relationships, and built-ins for literal search and focused reading.
Proposal
Acceptance criteria
Non-goals / out of scope
Edge cases
Context
navigate-code/SKILL.mdROADMAP.md