Skip to content

feat(generated-detect): suppress bug-fix attribution for auto-generated files - #19

Merged
avifenesh merged 2 commits into
mainfrom
feat/detect-generated-files
Apr 23, 2026
Merged

avifenesh merged 2 commits into
mainfrom
feat/detect-generated-files

Conversation

@avifenesh

@avifenesh avifenesh commented Apr 23, 2026 •

Copy link
Copy Markdown
Contributor

Why

Bug-fix attribution propagates badly through generated files. A fix(schema): correct field type commit touches both the .proto source and the mechanically-derived .pb.go bindings, so the bindings end up with a bug_fix_changes count even though they are artifacts of the schema fix. The bindings then float to the top of bugspots/painspots queries and drown out actually-broken code.

What changes

New analyzer-core::generated_detect module with one public function, is_generated_path(path), wired into the aggregator. When a file's path looks generated, its FileActivity entry is marked generated=true and bug_fix_changes is no longer incremented for that file (even when a fix-prefix commit touches it).

The other fields - changes, recent_changes, additions, deletions, authors, coupling - continue to be populated, so co-change graphs and ownership queries remain useful.

Detection heuristics (path-only, conservative)

  • Directory markers: /generated/, /__generated__/, /.generated/, /codegen/, /grpc-gen/, /autogenerated/, /auto-generated/
  • Protobuf/gRPC bindings: *.pb.go, *.pb.cc, *.pb.h, *_pb.py, *_pb2.py, *_pb2_grpc.py, *.pb.ts, *.pb.js, *_grpc.pb.go
  • Type bindings: *.d.ts
  • GraphQL/OpenAPI/Swagger: *.generated.*, *.openapi.*, *.swagger.*
  • Snapshot artifacts: *.snap
  • Stable basenames: generated.{rs,go,ts,js,py}, schema.generated.*

Conservative bias: when in doubt, treat as human-authored. Words like regenerate and degenerate do not match (substring scan only matches /generated/, with the slashes). The cost of a false negative (one extra bugspots row) is far lower than a false positive (silently losing bug signal on real source).

Out of scope (can layer on later)

  • @generated header marker detection
  • // Code generated by ... DO NOT EDIT. detection
  • .gitattributes linguist-generated=true parsing

These all require per-file content reads that the aggregator currently does not perform.

Schema change

FileActivity gains a generated: bool field with #[serde(default)], so older repo-intel artifacts deserialize cleanly (generated defaults to false, then gets refreshed on next init/update).

Tests

  • 8 new unit tests in generated_detect.rs covering positive cases (proto bindings, codegen dirs, type defs, snapshots, Windows paths) and negative cases (regenerate_token.rs, degenerate_input.rs, empty path, plain source files).
  • 1 new aggregator-level test confirming a fix(schema) commit credits the .proto source but not the .pb.go bindings or generated/ output.

Test plan

  • cargo test --workspace - all tests pass
  • cargo clippy --workspace --all-targets -- -D warnings - clean
  • cargo fmt --all - clean
  • CI green on PR

Independence

Stacks cleanly with #17 (drop-AI-detection) and #18 (broaden-bugfix-detect). All three PRs touch different code paths and either order works. They can land in any order.


Note

Medium Risk
Changes the repo-intel JSON schema (FileActivity.generated) and alters bug-fix counting logic, which will affect downstream bugspots/painspots results and any consumers expecting the old schema.

Overview
Adds path-based generated-file detection (generated_detect::is_generated_path) and records the result in a new FileActivity.generated field (serde-defaulted for backward compatibility).

Updates the git-map aggregator to set generated on first sight of a file and to skip incrementing bug_fix_changes for generated paths, reducing false bugspot/painspot signals from codegen artifacts; includes new unit tests for the detector and an aggregator test covering the suppression behavior.

Reviewed by Cursor Bugbot for commit 6522feb. Configure here.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a heuristic detection system for auto-generated files to prevent mechanical artifacts from inflating bug-fix metrics. It adds a new generated_detect module, updates the FileActivity data structure, and modifies the git aggregator to suppress bug-fix attribution for identified files. Review feedback suggests refining the path matching logic to correctly handle root-level directories, using path component splitting for more robust directory identification, and adopting RegexSet to improve the efficiency of pattern matching.

Comment on lines +39 to +50
const GENERATED_SUBSTRINGS: &[&str] = &[
// Conventional output directories and naming
"/generated/",
"/__generated__/",
"/.generated/",
"/autogenerated/",
"/auto-generated/",
"/codegen/",
// Stub/binding output directories
"/grpc-gen/",
"/grpc_gen/",
];

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The current GENERATED_SUBSTRINGS patterns (e.g., "/generated/") will fail to match generated directories located at the root of the repository because relative paths in git often lack a leading slash (e.g., generated/types.rs). While the slashes are intended to prevent false positives like regenerate_token.rs, they make the heuristic too conservative for top-level directories.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 05b8368 — switched to component-wise matching: split path on / and compare each component to the known generated dir names. Top-level generated/types.rs now matches, while regenerate_token.rs and degenerate_input.rs stay rejected. New test top_level_generated_directories_hit covers 5 root-level cases.

Comment on lines +111 to +115
for needle in GENERATED_SUBSTRINGS {
if normalized.contains(needle) {
return true;
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Instead of a substring scan with slashes, consider splitting the normalized path into components and checking if any component exactly matches a set of known generated directory names. This is more robust as it correctly identifies top-level directories while still avoiding partial matches like regenerate.

    // Check if any path component matches a known generated directory name.
    if normalized.split('/').any(|s| matches!(s, "generated" | "__generated__" | ".generated" | "autogenerated" | "auto-generated" | "codegen" | "grpc-gen" | "grpc_gen")) {
        return true;
    }

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 05b8368 — exactly this: for component in normalized.split('/') checked against GENERATED_DIR_COMPONENTS. Cleaner intent and correctly handles repo-root paths.

Comment on lines +123 to +127
for re in generated_basename_patterns() {
if re.is_match(basename) {
return true;
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Iterating through a list of regexes and calling is_match on each is less efficient than using a RegexSet. A RegexSet can match multiple patterns in a single scan of the input string.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 05b8368 — Vec<Regex> + per-pattern is_match loop replaced with a single RegexSet. One DFA, one scan.

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 2 potential issues.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 6522feb. Configure here.

bug_fix_changes: 0,
refactor_changes: 0,
last_bug_fix: String::new(),
generated: is_generated_path(&file.path),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Generated flag never refreshed for pre-existing file entries

Medium Severity

The generated flag is only set inside or_insert_with, which runs exclusively when a FileActivity entry is first created. During an incremental update on a map deserialized from before this feature, pre-existing entries retain generated: false (from #[serde(default)]) because or_insert_with is not invoked for them. This means bug-fix attribution is silently not suppressed for those generated files — exactly the scenario the feature aims to prevent. The flag needs to be unconditionally refreshed after the or_insert_with call.

Additional Locations (1)
Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 6522feb. Configure here.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 05b8368 — generated is now refreshed on every visit, not just on insert. Comment explains why (deserialized maps from before the field existed have generated: false from #[serde(default)] and would never get marked otherwise). Test test_merge_delta_refreshes_generated_flag_on_existing_entries pre-seeds an entry with generated: false and asserts the flag flips after merge_delta.

// Stub/binding output directories
"/grpc-gen/",
"/grpc_gen/",
];

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Directory markers miss top-level generated directories

Low Severity

Every entry in GENERATED_SUBSTRINGS starts with / (e.g., "/generated/"), and matching uses normalized.contains(needle). Git paths are repo-root-relative with no leading slash, so a file in a top-level generated directory like "generated/types.rs" does not contain "/generated/" and is silently missed. Only nested paths like "src/generated/types.rs" match. This affects all eight directory markers equally — codegen/, __generated__/, etc. at the repo root all produce false negatives.

Additional Locations (1)
Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 6522feb. Configure here.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 05b8368 — switched from substring scan with leading slashes to component-wise matching. Top-level generated/types.rs, codegen/foo.go, etc. now register correctly while substring collisions like regenerate_token.rs stay rejected.

@avifenesh
avifenesh force-pushed the feat/detect-generated-files branch from 05b8368 to 60c83d7 Compare April 23, 2026 22:46
…ed files

Background: bug-fix attribution propagates badly through generated
files. A "fix(schema): correct field type" commit touches both the
.proto source AND the mechanically-derived .pb.go bindings, so the
bindings end up with a bug_fix_changes count even though they are
artifacts of the schema fix. The bindings then float to the top of
bugspots/painspots queries and drown out actually-broken code.

This adds a new analyzer-core::generated_detect module with a single
public function, is_generated_path(path), and wires it into the
aggregator. When a file's path looks generated, its FileActivity entry
is marked generated=true and bug_fix_changes is no longer incremented
for that file (even when a fix-prefix commit touches it). The other
fields - changes, recent_changes, additions, deletions, authors,
coupling - continue to be populated, so co-change graphs and ownership
queries remain useful.

Detection heuristics (path-only, conservative):
- Directory markers: /generated/, /__generated__/, /.generated/,
  /codegen/, /grpc-gen/, /autogenerated/, /auto-generated/
- Protobuf/gRPC bindings: *.pb.go, *.pb.cc, *.pb.h, *_pb.py,
  *_pb2.py, *_pb2_grpc.py, *.pb.ts, *.pb.js, *_grpc.pb.go
- Type bindings: *.d.ts
- GraphQL/OpenAPI/Swagger: *.generated.*, *.openapi.*, *.swagger.*
- Snapshot artifacts: *.snap
- Stable basenames: generated.{rs,go,ts,js,py}, schema.generated.*

Conservative bias: when in doubt, treat as human-authored. Words like
"regenerate" and "degenerate" do NOT match (substring scan only matches
/generated/, with the slashes). The cost of a false negative (one
extra bugspots row) is far lower than a false positive (silently
losing bug signal on real source).

Out of scope (can layer on later):
- @generated header marker detection
- // Code generated by ... DO NOT EDIT. detection
- .gitattributes linguist-generated=true parsing
These all require per-file content reads that the aggregator currently
does not perform.

Schema change: FileActivity gains a `generated: bool` field with
(generated defaults to false, then gets refreshed on next init/update).

Tests:
- 8 new unit tests in generated_detect.rs covering positive cases
  (proto bindings, codegen dirs, type defs, snapshots, Windows paths)
  and negative cases (regenerate_token.rs, degenerate_input.rs, empty
  path, plain source files).
- 1 new aggregator-level test confirming a fix(schema) commit credits
  the .proto source but not the .pb.go bindings or generated/ output.

Tests: workspace passes, clippy clean.
Two reviewer-caught bugs and two suggestions, all addressed:

1. Top-level generated directories were silently missed (gemini, cursor)

   The previous substring scan keyed on `/generated/`, `/codegen/`,
   etc., but git paths are repo-root-relative and never carry a leading
   slash. So `generated/types.rs` at the repo root never matched, while
   the same path nested under `src/` did. Switched to component-wise
   matching: split on `/` and check whether any component equals one
   of the known generated dir names. This still rejects substring
   collisions like `regenerate_token.rs` and `degenerate_input.rs`.

2. `generated` flag never refreshed on existing entries (cursor)

   `or_insert_with` only ran on first insert. When merge_delta runs
   against a map deserialized from before the field existed, all
   pre-existing entries had generated=false (from #[serde(default)])
   and the flag stayed false forever - bug-fix attribution kept
   leaking through to generated files. Now we re-classify on every
   visit so existing entries pick up the flag on the next
   init/update cycle.

3. Suggestion: use RegexSet (gemini)

   Switched from `Vec<Regex>` + per-pattern `is_match` loop to a
   single `RegexSet`. Builds one DFA, scans the input once.

4. Suggestion: tests for top-level dirs

   Added top_level_generated_directories_hit covering 5 root-level
   variants (generated/, __generated__/, codegen/, autogenerated/,
   auto-generated/) and a refresh-flag-on-update aggregator test
   that pre-seeds an entry with generated=false to simulate the
   serde-default scenario.

Reviewer: gemini-code-assist#19 (3x medium), cursor#19 (medium + low)
@avifenesh
avifenesh force-pushed the feat/detect-generated-files branch from 60c83d7 to 182c87a Compare April 23, 2026 22:54
@avifenesh
avifenesh merged commit 5869941 into main Apr 23, 2026
4 checks passed
avifenesh added a commit that referenced this pull request Apr 24, 2026
Bump workspace version to 0.5.0. Includes 6 merged PRs since v0.4.0:
drop AI attribution detection (#17), broaden bug-fix classification
(#18), suppress generated-file bugspot pollution (#19), entry-points
query (#23), find query (#24), and LLM-augmented descriptors/summary
subcommands (#25).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant