Skip to content

feat(skill): add routing regression cases for navigate-code #7

Description

@vkuprin

Summary

Turn the navigate-code routing contract into regression cases so skill changes are driven by observed failures rather than by adding more prose.

The cases should verify which route is selected first, which follow-up evidence is required, and which claims are forbidden—for example, treating name-based text matches as semantic references.

Motivation

The skill is intentionally small. Without regression cases, future contributors have only manual judgment when deciding whether a new rule improves routing or bloats agent context.

This work protects the core product behavior: Semble for behavior-level discovery, Serena for symbols and relationships, and built-ins for literal search and focused reading.

Proposal

  • Define routing cases for unknown behavior, exact text, known-file reasoning, references, implementations, renames, and architecture tracing.
  • Score the first selected route, unnecessary calls, evidence quality, and prohibited claims separately from final-answer correctness.
  • Include Serena project-activation behavior, 0-based Serena line handling, and Semble content-scope selection.
  • Run cases against both Claude Code and Codex when automation permits; otherwise use a shared case format with client-specific runners.
  • Require a failing case before materially expanding the skill body.

Acceptance criteria

  • Every current routing rule has at least one positive case.
  • High-risk failures have negative cases, including name grep reported as a reference check.
  • Results distinguish routing, efficiency, evidence quality, and answer correctness.
  • The same case definitions can be used by Claude Code and Codex runners.
  • The suite reports skill context size so instruction growth is visible.
  • Contributor documentation explains when a routing rule may be added or removed.

Non-goals / out of scope

  • Testing every Serena or Semble tool independently.
  • Replacing repository-level correctness benchmarks.
  • Requiring identical tool-call sequences from different models.

Edge cases

  • A requested MCP tool is unavailable despite the plugin being installed.
  • A symbol is exported by both application code and a dependency.
  • Dynamic-language references cannot be proven by the language server.
  • The correct route changes after semantic discovery identifies a concrete symbol.

Context

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area: skillnavigate-code routing and agent guidancep1High priority — trust, safety, or release confidence✨ featureNew user-facing capability

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions