Repository navigation
feat(llm-augmented): set-descriptors + set-summary subcommands; descriptor-aware find (PR1 of 2) - #25
Merged
Merged
Conversation
…iptor-aware find
Adds the storage + read side of LLM-augmented signals. The Rust crate
itself never calls an LLM - the orchestrating skill (in agent-sh/repo-intel)
spawns Haiku subagents via the Task tool, then pipes the agent output
back through these new subcommands.
This is PR1 of a 2-part series:
- THIS PR (agent-analyzer): passive storage + scoring of LLM-generated
fields, plus query surface to read them back.
- PR2 (agent-sh/repo-intel): two Haiku agents (repo-intel-weighter,
repo-intel-summarizer) and the skill orchestration that calls into
these new Rust subcommands.
## New subcommands
```
repo-intel set-descriptors --map-file FILE [--input PATH | --input -]
repo-intel set-summary --map-file FILE [--input PATH | --input -]
```
Both accept JSON via file path or stdin. set-descriptors merges into
existing entries (partial updates safe). set-summary fully replaces.
Both bump map.updated.
Schema:
- descriptors input: `{"path": "1-2 sentence descriptor", ...}`
- summary input: `{depth1, depth3, depth10, inputHash}`
## New query
```
repo-intel query summary <repo> --map-file FILE [--depth 1|3|10]
```
`--depth` returns just that depth as plain text; omit it for full JSON.
## find now consumes descriptors
`find()` gains an optional `descriptors` param. When the artifact has
`fileDescriptors` populated, each query term that matches a file's
descriptor adds 2.5 to the score - between path-component (3.0) and
import (1.0). Unlike the doc-header pass, the descriptor signal fires
even on files with zero cheap signals - that's the synonym-recall case
(file named `executor.rs` should surface for query `worker` if the
descriptor calls it a "worker pool implementation").
## Schema (RepoIntelData)
Adds two new optional, serde(default) fields under "Phase 6":
- `file_descriptors: Option<HashMap<String, String>>`
- `summary: Option<RepoSummary>` where `RepoSummary { depth1, depth3,
depth10, inputHash, generatedAt }`
Both default to None on existing artifacts. Zero migration needed.
## Tests
- 6 new analyzer-cli tests (set-descriptors merge + partial-update +
malformed-input rejection; set-summary round-trip + overwrite;
read_input file path)
- 3 new find tests (descriptor surfacing semantic synonym, descriptor
scores stack with v0 signals, descriptor scoring is term-independent)
- All existing tests updated to pass `None` descriptors
- 213 total workspace tests pass; clippy clean
## End-to-end smoke on agnix
Pipe descriptor via stdin, then query find:
```
$ echo '{"crates/agnix-core/src/file_utils.rs": "Worker pool dispatcher — spawns N background threads..."}' \
| agent-analyzer repo-intel set-descriptors --map-file ... --input -
[OK] merged 1 descriptors; total now 1
$ agent-analyzer repo-intel query find "worker pool" ../agnix --map-file ... --top 1
[{"path": "crates/agnix-core/src/file_utils.rs", "score": 5.0,
"why": "descriptor matches"}]
```
The file has no symbol/path/basename match for "worker" - it surfaced
purely from the descriptor. That's the v1 use case working end-to-end.
Same for set-summary + query summary --depth 1.
## What's deliberately NOT here
- No HTTP client. The Rust crate is offline-only.
- No Anthropic SDK or async runtime. The LLM lives in the skill layer.
- No API key handling. Same reason.
- No agent files (.md). Those go in the repo-intel plugin (PR2).
- No automatic post-init invocation. The skill orchestrates that.
## What lands in PR2
- `repo-intel-weighter.md` agent (Haiku) — generates the descriptor
JSON this PR's set-descriptors expects
- `repo-intel-summarizer.md` agent (Haiku) — generates the summary JSON
- `commands/repo-intel.md` updates — spawn each agent post-init and
pipe their output to the new Rust subcommands
- `lib/repo-intel/index.js` — `applyDescriptors` / `applySummary` JS
wrappers calling the new Rust subcommands
|
Warning You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #22 (storage half) and #20 v1 (storage half). PR2 in agent-sh/repo-intel will land the Haiku agents and the skill orchestration that drives these new subcommands.
Why a 2-PR split
The Rust crate stays offline-only by design — it never calls an LLM. The summary and descriptor data come from Haiku subagents spawned by the JS skill via the Task tool. This PR adds the storage and read side so the skill can pipe the agents' output back through:
```
repo-intel set-descriptors --map-file FILE --input -
repo-intel set-summary --map-file FILE --input -
```
PR2 adds the agents themselves and the skill orchestration.
What's in this PR
New top-level subcommands
New `summary` query
```
repo-intel query summary --map-file FILE [--depth 1|3|10]
```
`--depth` returns just that depth as plain text; omit it for full JSON. Returns `null` with a hint when no summary is present.
`find` becomes descriptor-aware
`find()` gains an optional `descriptors` param. When the artifact has `fileDescriptors`, each query term that matches a file's descriptor adds 2.5 to the score — between path-component (3.0) and import (1.0).
Unlike the doc-header pass, the descriptor signal fires even on files with zero cheap signals — that's the synonym recall case (a file named `executor.rs` should surface for query `worker` if the descriptor says it's a worker pool).
Schema (RepoIntelData, "Phase 6")
Two new optional, `#[serde(default)]` fields:
Both default to `None` on existing artifacts. Zero migration.
End-to-end smoke on agnix
```
$ echo '{"crates/agnix-core/src/file_utils.rs": "Worker pool dispatcher — spawns N background threads, drains an mpsc queue on shutdown."}'
| agent-analyzer repo-intel set-descriptors --map-file ... --input -
[OK] merged 1 descriptors; total now 1
$ agent-analyzer repo-intel query find "worker pool" ../agnix --map-file ... --top 1
[{"path": "crates/agnix-core/src/file_utils.rs", "score": 5.0,
"why": "descriptor matches"}]
```
That file has zero substring/symbol/import match for "worker" — it surfaced purely from the descriptor. The v1 use case working end-to-end.
Same end-to-end works for `set-summary` + `query summary --depth 1`.
Tests
What's deliberately NOT here
All of that lives in the skill layer (PR2 in agent-sh/repo-intel).
What's in PR2 (next)
Test plan
Note
Medium Risk
Adds new persisted fields to the
RepoIntelDataartifact and changesfindranking behavior, so consumers/tests relying on the prior schema or scoring may break. Risk is moderated by the fields being optional with defaults and covered by new unit tests.Overview
Introduces a new Phase 6 extension to the repo-intel artifact: optional
file_descriptorsandsummary(RepoSummarywithgenerated_at) stored inRepoIntelDatawith backwards-compatible serde defaults.Extends the
agent-analyzer repo-intelCLI withset-descriptors(merge{path: descriptor}from file/stdin) andset-summary(replace{depth1, depth3, depth10, inputHash}from file/stdin), plus a newrepo-intel query summaryoutput mode.Updates
analyzer-repo-map’sfindscorer to accept optional descriptors and add a new descriptor-match signal (2.5/term), wiring this through the CLI and adding targeted tests; workspace fixtures that constructRepoIntelDataare updated to include the new fields.Reviewed by Cursor Bugbot for commit 3b987e3. Configure here.