Skip to content

feat(llm-augmented): set-descriptors + set-summary subcommands; descriptor-aware find (PR1 of 2) - #25

Merged
avifenesh merged 1 commit into
mainfrom
feat/haiku-summary-and-find-v1
Apr 24, 2026
Merged

avifenesh merged 1 commit into
mainfrom
feat/haiku-summary-and-find-v1

Conversation

@avifenesh

@avifenesh avifenesh commented Apr 24, 2026 •

Copy link
Copy Markdown
Contributor

Closes #22 (storage half) and #20 v1 (storage half). PR2 in agent-sh/repo-intel will land the Haiku agents and the skill orchestration that drives these new subcommands.

Why a 2-PR split

The Rust crate stays offline-only by design — it never calls an LLM. The summary and descriptor data come from Haiku subagents spawned by the JS skill via the Task tool. This PR adds the storage and read side so the skill can pipe the agents' output back through:

```
repo-intel set-descriptors --map-file FILE --input -
repo-intel set-summary --map-file FILE --input -
```

PR2 adds the agents themselves and the skill orchestration.

What's in this PR

New top-level subcommands

Command Input Behavior
`set-descriptors` JSON `{path: descriptor}` (file or stdin) Merges into existing entries (partial updates safe). Bumps `updated`.
`set-summary` JSON `{depth1, depth3, depth10, inputHash}` (file or stdin) Replaces the previous summary. Stamps `generatedAt`.

New `summary` query

```
repo-intel query summary --map-file FILE [--depth 1|3|10]
```

`--depth` returns just that depth as plain text; omit it for full JSON. Returns `null` with a hint when no summary is present.

`find` becomes descriptor-aware

`find()` gains an optional `descriptors` param. When the artifact has `fileDescriptors`, each query term that matches a file's descriptor adds 2.5 to the score — between path-component (3.0) and import (1.0).

Unlike the doc-header pass, the descriptor signal fires even on files with zero cheap signals — that's the synonym recall case (a file named `executor.rs` should surface for query `worker` if the descriptor says it's a worker pool).

Schema (RepoIntelData, "Phase 6")

Two new optional, `#[serde(default)]` fields:

  • `file_descriptors: Option<HashMap<String, String>>`
  • `summary: Option` where `RepoSummary { depth1, depth3, depth10, inputHash, generatedAt }`

Both default to `None` on existing artifacts. Zero migration.

End-to-end smoke on agnix

```
$ echo '{"crates/agnix-core/src/file_utils.rs": "Worker pool dispatcher — spawns N background threads, drains an mpsc queue on shutdown."}'
| agent-analyzer repo-intel set-descriptors --map-file ... --input -
[OK] merged 1 descriptors; total now 1

$ agent-analyzer repo-intel query find "worker pool" ../agnix --map-file ... --top 1
[{"path": "crates/agnix-core/src/file_utils.rs", "score": 5.0,
"why": "descriptor matches"}]
```

That file has zero substring/symbol/import match for "worker" — it surfaced purely from the descriptor. The v1 use case working end-to-end.

Same end-to-end works for `set-summary` + `query summary --depth 1`.

Tests

  • 6 new analyzer-cli tests (set-descriptors merge + partial-update + malformed-input rejection; set-summary round-trip + overwrite; read_input file-vs-stdin)
  • 3 new find tests (descriptor surfaces semantic synonym, scores stack with v0 signals, term-independent)
  • All existing tests updated to pass `None` descriptors
  • 213 total workspace tests pass; clippy clean

What's deliberately NOT here

  • No HTTP client / Anthropic SDK / async runtime
  • No API key handling
  • No agent .md files
  • No automatic post-init invocation

All of that lives in the skill layer (PR2 in agent-sh/repo-intel).

What's in PR2 (next)

  • `repo-intel-weighter.md` agent (Haiku) — generates descriptor JSON
  • `repo-intel-summarizer.md` agent (Haiku) — generates summary JSON
  • `commands/repo-intel.md` updates — spawn each agent post-init, pipe to these new subcommands
  • `lib/repo-intel/index.js` — `applyDescriptors` / `applySummary` JS wrappers

Test plan

  • `cargo test --workspace` — 213 pass
  • `cargo clippy --workspace --all-targets -- -D warnings` — clean
  • `cargo fmt --all` — clean
  • Smoke-tested both subcommands + descriptor-aware find end-to-end on agnix
  • CI green on PR

Note

Medium Risk
Adds new persisted fields to the RepoIntelData artifact and changes find ranking behavior, so consumers/tests relying on the prior schema or scoring may break. Risk is moderated by the fields being optional with defaults and covered by new unit tests.

Overview
Introduces a new Phase 6 extension to the repo-intel artifact: optional file_descriptors and summary (RepoSummary with generated_at) stored in RepoIntelData with backwards-compatible serde defaults.

Extends the agent-analyzer repo-intel CLI with set-descriptors (merge {path: descriptor} from file/stdin) and set-summary (replace {depth1, depth3, depth10, inputHash} from file/stdin), plus a new repo-intel query summary output mode.

Updates analyzer-repo-map’s find scorer to accept optional descriptors and add a new descriptor-match signal (2.5/term), wiring this through the CLI and adding targeted tests; workspace fixtures that construct RepoIntelData are updated to include the new fields.

Reviewed by Cursor Bugbot for commit 3b987e3. Configure here.

…iptor-aware find

Adds the storage + read side of LLM-augmented signals. The Rust crate
itself never calls an LLM - the orchestrating skill (in agent-sh/repo-intel)
spawns Haiku subagents via the Task tool, then pipes the agent output
back through these new subcommands.

This is PR1 of a 2-part series:
- THIS PR (agent-analyzer): passive storage + scoring of LLM-generated
  fields, plus query surface to read them back.
- PR2 (agent-sh/repo-intel): two Haiku agents (repo-intel-weighter,
  repo-intel-summarizer) and the skill orchestration that calls into
  these new Rust subcommands.

## New subcommands

```
repo-intel set-descriptors --map-file FILE [--input PATH | --input -]
repo-intel set-summary     --map-file FILE [--input PATH | --input -]
```

Both accept JSON via file path or stdin. set-descriptors merges into
existing entries (partial updates safe). set-summary fully replaces.
Both bump map.updated.

Schema:
- descriptors input: `{"path": "1-2 sentence descriptor", ...}`
- summary input:    `{depth1, depth3, depth10, inputHash}`

## New query

```
repo-intel query summary <repo> --map-file FILE [--depth 1|3|10]
```

`--depth` returns just that depth as plain text; omit it for full JSON.

## find now consumes descriptors

`find()` gains an optional `descriptors` param. When the artifact has
`fileDescriptors` populated, each query term that matches a file's
descriptor adds 2.5 to the score - between path-component (3.0) and
import (1.0). Unlike the doc-header pass, the descriptor signal fires
even on files with zero cheap signals - that's the synonym-recall case
(file named `executor.rs` should surface for query `worker` if the
descriptor calls it a "worker pool implementation").

## Schema (RepoIntelData)

Adds two new optional, serde(default) fields under "Phase 6":
- `file_descriptors: Option<HashMap<String, String>>`
- `summary: Option<RepoSummary>` where `RepoSummary { depth1, depth3,
  depth10, inputHash, generatedAt }`

Both default to None on existing artifacts. Zero migration needed.

## Tests

- 6 new analyzer-cli tests (set-descriptors merge + partial-update +
  malformed-input rejection; set-summary round-trip + overwrite;
  read_input file path)
- 3 new find tests (descriptor surfacing semantic synonym, descriptor
  scores stack with v0 signals, descriptor scoring is term-independent)
- All existing tests updated to pass `None` descriptors
- 213 total workspace tests pass; clippy clean

## End-to-end smoke on agnix

Pipe descriptor via stdin, then query find:
```
$ echo '{"crates/agnix-core/src/file_utils.rs": "Worker pool dispatcher — spawns N background threads..."}' \
    | agent-analyzer repo-intel set-descriptors --map-file ... --input -
[OK] merged 1 descriptors; total now 1

$ agent-analyzer repo-intel query find "worker pool" ../agnix --map-file ... --top 1
[{"path": "crates/agnix-core/src/file_utils.rs", "score": 5.0,
  "why": "descriptor matches"}]
```

The file has no symbol/path/basename match for "worker" - it surfaced
purely from the descriptor. That's the v1 use case working end-to-end.

Same for set-summary + query summary --depth 1.

## What's deliberately NOT here

- No HTTP client. The Rust crate is offline-only.
- No Anthropic SDK or async runtime. The LLM lives in the skill layer.
- No API key handling. Same reason.
- No agent files (.md). Those go in the repo-intel plugin (PR2).
- No automatic post-init invocation. The skill orchestrates that.

## What lands in PR2

- `repo-intel-weighter.md` agent (Haiku) — generates the descriptor
  JSON this PR's set-descriptors expects
- `repo-intel-summarizer.md` agent (Haiku) — generates the summary JSON
- `commands/repo-intel.md` updates — spawn each agent post-init and
  pipe their output to the new Rust subcommands
- `lib/repo-intel/index.js` — `applyDescriptors` / `applySummary` JS
  wrappers calling the new Rust subcommands
@gemini-code-assist

Copy link
Copy Markdown

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@avifenesh
avifenesh merged commit 55eaeb5 into main Apr 24, 2026
5 checks passed
avifenesh added a commit that referenced this pull request Apr 24, 2026
Bump workspace version to 0.5.0. Includes 6 merged PRs since v0.4.0:
drop AI attribution detection (#17), broaden bug-fix classification
(#18), suppress generated-file bugspot pollution (#19), entry-points
query (#23), find query (#24), and LLM-augmented descriptors/summary
subcommands (#25).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat: summary [--depth 1|3|10] query — narrative repo description, generated with Haiku at init/update

1 participant