Repository navigation
Stop hook: verify thinking quality at session end — task completeness, assumptions, stale logs, disk space (delivery-gate) - #2378
Conversation
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
📝 WalkthroughSummary by CodeRabbit
WalkthroughAdds a new ChangesDelivery Gate Stop-Hook
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~25 minutes Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
| # 3. Rationalization pattern detection | ||
| hits = [] | ||
| for p in RATIONALIZE: | ||
| m = re.search(p, tail, re.IGNORECASE) | ||
| if m: | ||
| hits.append(m.group(0)[:80]) | ||
| if hits: | ||
| log.warning('quality-gate: %s', hits) |
There was a problem hiding this comment.
Rationalization detection logs but never blocks
The PR description lists rationalization pattern detection as one of the two gate conditions, but hits only produces a log.warning and has no path to sys.exit(2). A session where Claude explicitly writes "skipping tests for now" passes through without being blocked. If this is intentionally advisory-only, the SKILL.md description should clarify that it warns but does not block, so operators set expectations correctly.
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
There was a problem hiding this comment.
Actionable comments posted: 5
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@skills/delivery-gate/hooks/quality-gate.py`:
- Around line 59-61: The project-memory key generation in quality-gate.py is
collision-prone because the manual cwd replacement in the memory path can map
different repos to the same directory. Replace the ad hoc safe variable logic
with the exact Claude project-path encoding or another collision-resistant
mapping in the hook that builds mem, so the memory lookup stays scoped to the
correct project and cannot read another repo’s memory.
- Around line 47-50: The reminder path in quality-gate.py is unreachable because
`logging.basicConfig(... level=logging.WARNING)` suppresses the `log.info(...)`
branch used by `DISK_REMIND_GB`. Update the logging setup or the `disk`
threshold handling in the quality gate logic so the reminder message can
actually be emitted, and make sure the same fix is applied in the related
reminder/warn/block flow referenced by the `DISK_REMIND_GB` checks.
- Around line 122-127: The count_edits helper is only scanning the last 8000
characters, so earlier Edit/Write tool calls can be missed and the quality gate
may undercount edits. Update count_edits in quality-gate.py to inspect the full
transcript (or a robust bounded window that cannot drop older tool calls) while
still matching the structured tool-call pattern used by the current regex. Keep
the counting logic centralized in count_edits so edit_count always reflects all
Edit/Write invocations before comparing against COMPLEX_THRESHOLD.
- Around line 162-169: The compliance check in quality-gate.py currently fails
open when get_project_memory_dir() returns None by setting stale to an empty
list, so complex edits can pass without any library validation. Update the
gating logic around get_project_memory_dir(), check_stale_libs(), and the stale
handling to fail closed: either raise a setup error when memory is unavailable
or treat the required libraries as stale so enforcement still blocks the task.
In `@skills/delivery-gate/SKILL.md`:
- Line 3: The current description for delivery-gate overstates what the hook
checks. Update the wording in SKILL.md so it matches the actual implementation
in the delivery-gate hook, which only enforces the fixed RATIONALIZE regex list,
edit/write counts, library mtimes, and disk space. Keep the summary accurate and
avoid claiming generic detection of contradictions, omissions, or unverified
assumptions unless the hook actually implements those checks.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: 49d6e8ed-6728-4c8d-ab85-5c38f400d9a6
📒 Files selected for processing (2)
skills/delivery-gate/SKILL.mdskills/delivery-gate/hooks/quality-gate.py
📜 Review details
⏰ Context from checks skipped due to timeout. (1)
- GitHub Check: Greptile Review
🧰 Additional context used
📓 Path-based instructions (13)
skills/**/*.md
📄 CodeRabbit inference engine (CLAUDE.md)
Skills should be formatted as Markdown with clear sections for When to Use, How It Works, and Examples.
Files:
skills/delivery-gate/SKILL.md
{agents,skills,commands}/**/*.md
📄 CodeRabbit inference engine (CLAUDE.md)
Use lowercase filenames with hyphens (e.g.,
python-reviewer.md,tdd-workflow.md) for agents, skills, and commands.
Files:
skills/delivery-gate/SKILL.md
{skills,commands,agents,rules}/**
⚙️ CodeRabbit configuration file
{skills,commands,agents,rules}/**: Focus on prompt-injection resilience, tool-permission scope, destructive action guards, and secret exfiltration risks.
Files:
skills/delivery-gate/SKILL.mdskills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,go,rb,php,scala,kt}
📄 CodeRabbit inference engine (.cursor/rules/common-coding-style.md)
**/*.{js,ts,jsx,tsx,py,java,cs,go,rb,php,scala,kt}: Always create new objects, never mutate existing ones. Use immutable patterns to prevent hidden side effects and enable safe concurrency
Organize code into many small files (200-400 lines typical, 800 lines max) organized by feature/domain rather than by type
Always handle errors explicitly at every level and never silently swallow errors
Always validate all user input before processing at system boundaries
Use schema-based validation where available
Fail fast with clear error messages when validation fails
Never trust external data (API responses, user input, file content)
Ensure code is readable and well-named
Keep functions small (less than 50 lines)
Keep files focused (less than 800 lines)
Avoid deep nesting (more than 4 levels)
Do not use hardcoded values; use constants or configuration instead
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php,swift,kt,rs,c,cpp,h,hpp}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
No hardcoded secrets (API keys, passwords, tokens) - validate before any commit
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php}: All user inputs must be validated
Enable CSRF protection on all state-changing endpoints
Verify authentication and authorization for all protected endpoints
Implement rate limiting on all endpoints to prevent abuse
Ensure error messages do not leak sensitive data in responses
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php,sql}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
Use parameterized queries to prevent SQL injection
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php,swift,kt,rs,c,cpp,h,hpp,properties,yml,yaml,json,env,config}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
NEVER hardcode secrets in source code - ALWAYS use environment variables or a secret manager
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{py,pyi}
📄 CodeRabbit inference engine (.cursor/rules/python-coding-style.md)
**/*.{py,pyi}: Follow PEP 8 conventions in Python code
Use type annotations on all function signatures in Python
Prefer immutable data structures such as frozen dataclasses and NamedTuple in Python
**/*.{py,pyi}: Auto-format Python files using black/ruff after edit
Run type checking using mypy/pyright after editing Python files
**/*.{py,pyi}: Use Protocol from typing module for duck typing and defining object shapes in Python
Use dataclasses with@dataclassdecorator for DTOs (Data Transfer Objects) in Python
Use context managers (with statement) for resource management in Python
Use generators for lazy evaluation and memory-efficient iteration in Python
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.py
📄 CodeRabbit inference engine (.cursor/rules/python-coding-style.md)
**/*.py: Use black for code formatting in Python
Use isort for import sorting in Python
Use ruff for linting Python codeAvoid using
print()statements in Python code; use theloggingmodule instead
**/*.py: Retrieve secrets and API keys from environment variables using os.environ with error handling (raise KeyError if missing) rather than hardcoding credentials
Use bandit for static security analysis in Python projects
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt,cpp,c,fs}
📄 CodeRabbit inference engine (AGENTS.md)
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt,cpp,c,fs}: Write tests before implementation using TDD workflow: write failing test (RED), implement minimal code (GREEN), then refactor (IMPROVE)
Keep functions small (<50 lines) and files focused (<800 lines, typical 200-400 lines)
Avoid deep nesting (>4 levels)
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt}
📄 CodeRabbit inference engine (AGENTS.md)
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt}: Never mutate existing objects; always create new objects with changes applied (Immutability requirement)
Handle errors at every level; provide user-friendly messages in UI code and detailed context in server-side logs
Ensure error messages don't leak sensitive data
Files:
skills/delivery-gate/hooks/quality-gate.py
skills/**/*.{js,ts,py,sh}
📄 CodeRabbit inference engine (AGENTS.md)
Place new workflow contributions in
skills/directory as the canonical workflow surface
Files:
skills/delivery-gate/hooks/quality-gate.py
🧠 Learnings (2)
📚 Learning: 2026-03-15T19:02:43.245Z
Learnt from: imrobinsingh
Repo: affaan-m/everything-claude-code PR: 503
File: skills/data-scraper-agent/SKILL.md:1-748
Timestamp: 2026-03-15T19:02:43.245Z
Learning: In this repository, skill folders should use a lowercase-hyphen name (e.g., data-scraper-agent, claude-api) and the skill description file inside each folder should be named SKILL.md (uppercase). Do not flag SKILL.md as a naming violation; treat SKILL.md as the canonical file name inside each skill directory.
Applied to files:
skills/delivery-gate/SKILL.md
📚 Learning: 2026-04-15T15:52:59.963Z
Learnt from: manja316
Repo: affaan-m/everything-claude-code PR: 1360
File: skills/security-bounty-hunter/SKILL.md:11-18
Timestamp: 2026-04-15T15:52:59.963Z
Learning: In this repository’s skills documentation (skills/**/SKILL.md), use the canonical auto-activation skill section header `## When to Activate`—do not use `## When to Use`. CONTRIBUTING.md and docs/SKILL-DEVELOPMENT-GUIDE.md confirm the required header, and existing skills follow this convention. This header is important for the auto-activation mechanism to detect the correct section.
Applied to files:
skills/delivery-gate/SKILL.md
🪛 ast-grep (0.44.0)
skills/delivery-gate/hooks/quality-gate.py
[warning] 154-154: Regex pattern passed to re is built from a non-literal (variable, call, concatenation, or f-string) value. If that value is attacker-controlled it can introduce a malicious pattern with catastrophic backtracking (ReDoS). Use a hardcoded literal pattern, or validate/escape untrusted input with re.escape() and bound the regex complexity before compiling.
Context: re.search(p, tail, re.IGNORECASE)
Note: [CWE-1333] Inefficient Regular Expression Complexity.
(redos-non-literal-regex-python)
🪛 Ruff (0.15.18)
skills/delivery-gate/hooks/quality-gate.py
[warning] 75-75: Consider moving this statement to an else block
(TRY300)
[warning] 75-75: Unnecessary assignment to free_gb before return statement
Remove unnecessary assignment
(RET504)
[warning] 82-82: Too many branches (13 > 12)
(PLR0912)
[warning] 86-86: datetime.date.today() used
(DTZ011)
[warning] 97-97: datetime.datetime.fromtimestamp() called without a tz argument
(DTZ006)
[warning] 109-109: datetime.datetime.fromtimestamp() called without a tz argument
(DTZ006)
[warning] 130-130: Too many branches (16 > 12)
(PLR0912)
[warning] 166-169: Use ternary operator stale = check_stale_libs(mem_dir) if mem_dir else [] instead of if-else-block
Replace if-else-block with stale = check_stale_libs(mem_dir) if mem_dir else []
(SIM108)
[warning] 176-176: zip() without an explicit strict= parameter
Add explicit value for parameter strict=
(B905)
[warning] 188-188: Logging statement uses f-string
(G004)
| cwd = os.environ.get('CLAUDE_PROJECT_DIR', os.getcwd()) | ||
| safe = cwd.replace(':', '').replace('\\', '-').replace('/', '-') | ||
| mem = os.path.expanduser(f'~/.claude/projects/{safe}/memory') |
There was a problem hiding this comment.
🔒 Security & Privacy | 🟠 Major | 🏗️ Heavy lift
Project-memory lookup is collision-prone.
cwd.replace(...) is lossy: /work/a-b/c and /work/a/b-c both normalize to the same key. That lets this hook read another project's memory and make allow/block decisions from the wrong repo. Use the exact Claude project-path encoding or a collision-resistant mapping instead of manual character replacement. As per path instructions, skills/** reviews should focus on secret exfiltration risks and permission scope.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@skills/delivery-gate/hooks/quality-gate.py` around lines 59 - 61, The
project-memory key generation in quality-gate.py is collision-prone because the
manual cwd replacement in the memory path can map different repos to the same
directory. Replace the ad hoc safe variable logic with the exact Claude
project-path encoding or another collision-resistant mapping in the hook that
builds mem, so the memory lookup stays scoped to the correct project and cannot
read another repo’s memory.
Source: Path instructions
There was a problem hiding this comment.
Actionable comments posted: 1
♻️ Duplicate comments (1)
skills/delivery-gate/hooks/quality-gate.py (1)
167-171: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick winMissing memory setup still disables the gate for complex tasks.
This branch warns, then sets
stale = [], so both block conditions at Lines 189-196 are skipped even though no learning library was checked. That bypasses the PR’s advertised “complex task without learning capture blocks stop” behavior. Treat a missing memory dir as stale for complex runs, or explicitly scope the feature as opt-in.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@skills/delivery-gate/hooks/quality-gate.py` around lines 167 - 171, The missing-memory branch in quality-gate.py currently warns and then clears the stale check, which lets complex tasks bypass enforcement. Update the logic in the memory-dir handling path so that when is_complex is true and no project memory directory is found, the run is treated as stale (or otherwise blocked) instead of setting stale to an empty list. Keep the change localized to the branch that logs the missing memory directory and the later stale/blocking conditions so the “complex task without learning capture blocks stop” behavior is preserved.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@skills/delivery-gate/hooks/quality-gate.py`:
- Around line 150-153: The rationalization pattern check is scanning
raw[-8000:], which can catch quoted user text or other Stop-payload fields
instead of the model’s actual reply. Update the logic in the quality-gate flow
around the RATIONALIZE loop to parse the payload and inspect only the last
assistant message field directly, so the regexes run against the assistant’s
final response rather than the raw tail.
---
Duplicate comments:
In `@skills/delivery-gate/hooks/quality-gate.py`:
- Around line 167-171: The missing-memory branch in quality-gate.py currently
warns and then clears the stale check, which lets complex tasks bypass
enforcement. Update the logic in the memory-dir handling path so that when
is_complex is true and no project memory directory is found, the run is treated
as stale (or otherwise blocked) instead of setting stale to an empty list. Keep
the change localized to the branch that logs the missing memory directory and
the later stale/blocking conditions so the “complex task without learning
capture blocks stop” behavior is preserved.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: b49b01dc-778c-4b31-9003-6714afe1e9dd
📒 Files selected for processing (2)
skills/delivery-gate/SKILL.mdskills/delivery-gate/hooks/quality-gate.py
📜 Review details
⏰ Context from checks skipped due to timeout. (1)
- GitHub Check: Greptile Review
🧰 Additional context used
📓 Path-based instructions (13)
skills/**/*.md
📄 CodeRabbit inference engine (CLAUDE.md)
Skills should be formatted as Markdown with clear sections for When to Use, How It Works, and Examples.
Files:
skills/delivery-gate/SKILL.md
{agents,skills,commands}/**/*.md
📄 CodeRabbit inference engine (CLAUDE.md)
Use lowercase filenames with hyphens (e.g.,
python-reviewer.md,tdd-workflow.md) for agents, skills, and commands.
Files:
skills/delivery-gate/SKILL.md
{skills,commands,agents,rules}/**
⚙️ CodeRabbit configuration file
{skills,commands,agents,rules}/**: Focus on prompt-injection resilience, tool-permission scope, destructive action guards, and secret exfiltration risks.
Files:
skills/delivery-gate/SKILL.mdskills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,go,rb,php,scala,kt}
📄 CodeRabbit inference engine (.cursor/rules/common-coding-style.md)
**/*.{js,ts,jsx,tsx,py,java,cs,go,rb,php,scala,kt}: Always create new objects, never mutate existing ones. Use immutable patterns to prevent hidden side effects and enable safe concurrency
Organize code into many small files (200-400 lines typical, 800 lines max) organized by feature/domain rather than by type
Always handle errors explicitly at every level and never silently swallow errors
Always validate all user input before processing at system boundaries
Use schema-based validation where available
Fail fast with clear error messages when validation fails
Never trust external data (API responses, user input, file content)
Ensure code is readable and well-named
Keep functions small (less than 50 lines)
Keep files focused (less than 800 lines)
Avoid deep nesting (more than 4 levels)
Do not use hardcoded values; use constants or configuration instead
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php,swift,kt,rs,c,cpp,h,hpp}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
No hardcoded secrets (API keys, passwords, tokens) - validate before any commit
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php}: All user inputs must be validated
Enable CSRF protection on all state-changing endpoints
Verify authentication and authorization for all protected endpoints
Implement rate limiting on all endpoints to prevent abuse
Ensure error messages do not leak sensitive data in responses
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php,sql}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
Use parameterized queries to prevent SQL injection
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php,swift,kt,rs,c,cpp,h,hpp,properties,yml,yaml,json,env,config}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
NEVER hardcode secrets in source code - ALWAYS use environment variables or a secret manager
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{py,pyi}
📄 CodeRabbit inference engine (.cursor/rules/python-coding-style.md)
**/*.{py,pyi}: Follow PEP 8 conventions in Python code
Use type annotations on all function signatures in Python
Prefer immutable data structures such as frozen dataclasses and NamedTuple in Python
**/*.{py,pyi}: Auto-format Python files using black/ruff after edit
Run type checking using mypy/pyright after editing Python files
**/*.{py,pyi}: Use Protocol from typing module for duck typing and defining object shapes in Python
Use dataclasses with@dataclassdecorator for DTOs (Data Transfer Objects) in Python
Use context managers (with statement) for resource management in Python
Use generators for lazy evaluation and memory-efficient iteration in Python
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.py
📄 CodeRabbit inference engine (.cursor/rules/python-coding-style.md)
**/*.py: Use black for code formatting in Python
Use isort for import sorting in Python
Use ruff for linting Python codeAvoid using
print()statements in Python code; use theloggingmodule instead
**/*.py: Retrieve secrets and API keys from environment variables using os.environ with error handling (raise KeyError if missing) rather than hardcoding credentials
Use bandit for static security analysis in Python projects
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt,cpp,c,fs}
📄 CodeRabbit inference engine (AGENTS.md)
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt,cpp,c,fs}: Write tests before implementation using TDD workflow: write failing test (RED), implement minimal code (GREEN), then refactor (IMPROVE)
Keep functions small (<50 lines) and files focused (<800 lines, typical 200-400 lines)
Avoid deep nesting (>4 levels)
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt}
📄 CodeRabbit inference engine (AGENTS.md)
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt}: Never mutate existing objects; always create new objects with changes applied (Immutability requirement)
Handle errors at every level; provide user-friendly messages in UI code and detailed context in server-side logs
Ensure error messages don't leak sensitive data
Files:
skills/delivery-gate/hooks/quality-gate.py
skills/**/*.{js,ts,py,sh}
📄 CodeRabbit inference engine (AGENTS.md)
Place new workflow contributions in
skills/directory as the canonical workflow surface
Files:
skills/delivery-gate/hooks/quality-gate.py
🧠 Learnings (2)
📚 Learning: 2026-03-15T19:02:43.245Z
Learnt from: imrobinsingh
Repo: affaan-m/everything-claude-code PR: 503
File: skills/data-scraper-agent/SKILL.md:1-748
Timestamp: 2026-03-15T19:02:43.245Z
Learning: In this repository, skill folders should use a lowercase-hyphen name (e.g., data-scraper-agent, claude-api) and the skill description file inside each folder should be named SKILL.md (uppercase). Do not flag SKILL.md as a naming violation; treat SKILL.md as the canonical file name inside each skill directory.
Applied to files:
skills/delivery-gate/SKILL.md
📚 Learning: 2026-04-15T15:52:59.963Z
Learnt from: manja316
Repo: affaan-m/everything-claude-code PR: 1360
File: skills/security-bounty-hunter/SKILL.md:11-18
Timestamp: 2026-04-15T15:52:59.963Z
Learning: In this repository’s skills documentation (skills/**/SKILL.md), use the canonical auto-activation skill section header `## When to Activate`—do not use `## When to Use`. CONTRIBUTING.md and docs/SKILL-DEVELOPMENT-GUIDE.md confirm the required header, and existing skills follow this convention. This header is important for the auto-activation mechanism to detect the correct section.
Applied to files:
skills/delivery-gate/SKILL.md
🪛 ast-grep (0.44.0)
skills/delivery-gate/hooks/quality-gate.py
[warning] 152-152: Regex pattern passed to re is built from a non-literal (variable, call, concatenation, or f-string) value. If that value is attacker-controlled it can introduce a malicious pattern with catastrophic backtracking (ReDoS). Use a hardcoded literal pattern, or validate/escape untrusted input with re.escape() and bound the regex complexity before compiling.
Context: re.search(p, raw[-8000:], re.IGNORECASE)
Note: [CWE-1333] Inefficient Regular Expression Complexity.
(redos-non-literal-regex-python)
🔇 Additional comments (1)
skills/delivery-gate/hooks/quality-gate.py (1)
50-50: LGTM!Also applies to: 123-127
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@skills/delivery-gate/hooks/quality-gate.py`:
- Around line 141-150: The transcript loading logic in the stdin parsing block
should not trust `transcript_path` from the JSON payload or silently ignore read
failures. In the `quality-gate.py` parsing path, validate that
`payload['transcript_path']` is a string, normalize and restrict it to the
expected Claude transcript root before using
`os.path.expanduser`/`os.path.exists`, and treat invalid or unreadable paths as
explicit failures instead of falling back via the broad `except` that currently
swallows `TypeError` and `OSError`. Keep the fallback to raw stdin only for
non-JSON input, and make the `transcript_path` branch fail closed when the file
cannot be safely read.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: f0a1dfef-31c9-4384-8feb-cb2f0ce2f04e
📒 Files selected for processing (1)
skills/delivery-gate/hooks/quality-gate.py
📜 Review details
⏰ Context from checks skipped due to timeout. (1)
- GitHub Check: Greptile Review
🧰 Additional context used
📓 Path-based instructions (11)
**/*.{js,ts,jsx,tsx,py,java,cs,go,rb,php,scala,kt}
📄 CodeRabbit inference engine (.cursor/rules/common-coding-style.md)
**/*.{js,ts,jsx,tsx,py,java,cs,go,rb,php,scala,kt}: Always create new objects, never mutate existing ones. Use immutable patterns to prevent hidden side effects and enable safe concurrency
Organize code into many small files (200-400 lines typical, 800 lines max) organized by feature/domain rather than by type
Always handle errors explicitly at every level and never silently swallow errors
Always validate all user input before processing at system boundaries
Use schema-based validation where available
Fail fast with clear error messages when validation fails
Never trust external data (API responses, user input, file content)
Ensure code is readable and well-named
Keep functions small (less than 50 lines)
Keep files focused (less than 800 lines)
Avoid deep nesting (more than 4 levels)
Do not use hardcoded values; use constants or configuration instead
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php,swift,kt,rs,c,cpp,h,hpp}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
No hardcoded secrets (API keys, passwords, tokens) - validate before any commit
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php}: All user inputs must be validated
Enable CSRF protection on all state-changing endpoints
Verify authentication and authorization for all protected endpoints
Implement rate limiting on all endpoints to prevent abuse
Ensure error messages do not leak sensitive data in responses
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php,sql}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
Use parameterized queries to prevent SQL injection
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php,swift,kt,rs,c,cpp,h,hpp,properties,yml,yaml,json,env,config}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
NEVER hardcode secrets in source code - ALWAYS use environment variables or a secret manager
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{py,pyi}
📄 CodeRabbit inference engine (.cursor/rules/python-coding-style.md)
**/*.{py,pyi}: Follow PEP 8 conventions in Python code
Use type annotations on all function signatures in Python
Prefer immutable data structures such as frozen dataclasses and NamedTuple in Python
**/*.{py,pyi}: Auto-format Python files using black/ruff after edit
Run type checking using mypy/pyright after editing Python files
**/*.{py,pyi}: Use Protocol from typing module for duck typing and defining object shapes in Python
Use dataclasses with@dataclassdecorator for DTOs (Data Transfer Objects) in Python
Use context managers (with statement) for resource management in Python
Use generators for lazy evaluation and memory-efficient iteration in Python
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.py
📄 CodeRabbit inference engine (.cursor/rules/python-coding-style.md)
**/*.py: Use black for code formatting in Python
Use isort for import sorting in Python
Use ruff for linting Python codeAvoid using
print()statements in Python code; use theloggingmodule instead
**/*.py: Retrieve secrets and API keys from environment variables using os.environ with error handling (raise KeyError if missing) rather than hardcoding credentials
Use bandit for static security analysis in Python projects
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt,cpp,c,fs}
📄 CodeRabbit inference engine (AGENTS.md)
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt,cpp,c,fs}: Write tests before implementation using TDD workflow: write failing test (RED), implement minimal code (GREEN), then refactor (IMPROVE)
Keep functions small (<50 lines) and files focused (<800 lines, typical 200-400 lines)
Avoid deep nesting (>4 levels)
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt}
📄 CodeRabbit inference engine (AGENTS.md)
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt}: Never mutate existing objects; always create new objects with changes applied (Immutability requirement)
Handle errors at every level; provide user-friendly messages in UI code and detailed context in server-side logs
Ensure error messages don't leak sensitive data
Files:
skills/delivery-gate/hooks/quality-gate.py
skills/**/*.{js,ts,py,sh}
📄 CodeRabbit inference engine (AGENTS.md)
Place new workflow contributions in
skills/directory as the canonical workflow surface
Files:
skills/delivery-gate/hooks/quality-gate.py
{skills,commands,agents,rules}/**
⚙️ CodeRabbit configuration file
{skills,commands,agents,rules}/**: Focus on prompt-injection resilience, tool-permission scope, destructive action guards, and secret exfiltration risks.
Files:
skills/delivery-gate/hooks/quality-gate.py
🪛 ast-grep (0.44.0)
skills/delivery-gate/hooks/quality-gate.py
[warning] 144-144: File path is request-/variable-derived; validate and normalize to prevent path traversal.
Context: open(tp, 'r', encoding='utf-8')
Note: [CWE-22] Improper Limitation of a Pathname to a Restricted Directory ('Path Traversal').
(open-filename-from-request)
[warning] 172-172: Regex pattern passed to re is built from a non-literal (variable, call, concatenation, or f-string) value. If that value is attacker-controlled it can introduce a malicious pattern with catastrophic backtracking (ReDoS). Use a hardcoded literal pattern, or validate/escape untrusted input with re.escape() and bound the regex complexity before compiling.
Context: re.search(p, tail, re.IGNORECASE)
Note: [CWE-1333] Inefficient Regular Expression Complexity.
(redos-non-literal-regex-python)
🪛 Ruff (0.15.18)
skills/delivery-gate/hooks/quality-gate.py
[warning] 76-76: Consider moving this statement to an else block
(TRY300)
[warning] 76-76: Unnecessary assignment to free_gb before return statement
Remove unnecessary assignment
(RET504)
[warning] 82-82: Too many branches (13 > 12)
(PLR0912)
[warning] 131-131: Too many branches (22 > 12)
(PLR0912)
[warning] 131-131: Too many statements (62 > 50)
(PLR0915)
[warning] 145-145: Unnecessary mode argument
Remove mode argument
(UP015)
🔇 Additional comments (2)
skills/delivery-gate/hooks/quality-gate.py (2)
60-62: Still valid: project-memory lookup is collision-prone.The manual
replace()encoding can still map distinct repos to the same~/.claude/projects/...directory, so this hook can read the wrong project’s memory and make allow/block decisions from another repo’s state.
168-177: Still valid: rationalization detection is scanning the transcript tail, not the last assistant reply.This still inspects
transcript[-8000:], so quoted user text or earlier messages can trigger warnings even when the final assistant response is clean.
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@skills/delivery-gate/hooks/quality-gate.py`:
- Line 60: The INFO log in quality-gate.py currently prints the full cwd and
resolved memory path, which can leak local project details in normal hook
output. Update the logging in the quality-gate hook to avoid exposing full paths
at INFO by either redacting the path values or moving this message to debug, and
keep the change localized to the hook’s logging call that uses log.info for the
memory-dir lookup.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: 8ef9a95c-2e8d-4af7-a276-452b4c1eb17b
📒 Files selected for processing (1)
skills/delivery-gate/hooks/quality-gate.py
📜 Review details
🧰 Additional context used
📓 Path-based instructions (11)
**/*.{js,ts,jsx,tsx,py,java,cs,go,rb,php,scala,kt}
📄 CodeRabbit inference engine (.cursor/rules/common-coding-style.md)
**/*.{js,ts,jsx,tsx,py,java,cs,go,rb,php,scala,kt}: Always create new objects, never mutate existing ones. Use immutable patterns to prevent hidden side effects and enable safe concurrency
Organize code into many small files (200-400 lines typical, 800 lines max) organized by feature/domain rather than by type
Always handle errors explicitly at every level and never silently swallow errors
Always validate all user input before processing at system boundaries
Use schema-based validation where available
Fail fast with clear error messages when validation fails
Never trust external data (API responses, user input, file content)
Ensure code is readable and well-named
Keep functions small (less than 50 lines)
Keep files focused (less than 800 lines)
Avoid deep nesting (more than 4 levels)
Do not use hardcoded values; use constants or configuration instead
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php,swift,kt,rs,c,cpp,h,hpp}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
No hardcoded secrets (API keys, passwords, tokens) - validate before any commit
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php}: All user inputs must be validated
Enable CSRF protection on all state-changing endpoints
Verify authentication and authorization for all protected endpoints
Implement rate limiting on all endpoints to prevent abuse
Ensure error messages do not leak sensitive data in responses
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php,sql}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
Use parameterized queries to prevent SQL injection
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php,swift,kt,rs,c,cpp,h,hpp,properties,yml,yaml,json,env,config}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
NEVER hardcode secrets in source code - ALWAYS use environment variables or a secret manager
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{py,pyi}
📄 CodeRabbit inference engine (.cursor/rules/python-coding-style.md)
**/*.{py,pyi}: Follow PEP 8 conventions in Python code
Use type annotations on all function signatures in Python
Prefer immutable data structures such as frozen dataclasses and NamedTuple in Python
**/*.{py,pyi}: Auto-format Python files using black/ruff after edit
Run type checking using mypy/pyright after editing Python files
**/*.{py,pyi}: Use Protocol from typing module for duck typing and defining object shapes in Python
Use dataclasses with@dataclassdecorator for DTOs (Data Transfer Objects) in Python
Use context managers (with statement) for resource management in Python
Use generators for lazy evaluation and memory-efficient iteration in Python
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.py
📄 CodeRabbit inference engine (.cursor/rules/python-coding-style.md)
**/*.py: Use black for code formatting in Python
Use isort for import sorting in Python
Use ruff for linting Python codeAvoid using
print()statements in Python code; use theloggingmodule instead
**/*.py: Retrieve secrets and API keys from environment variables using os.environ with error handling (raise KeyError if missing) rather than hardcoding credentials
Use bandit for static security analysis in Python projects
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt,cpp,c,fs}
📄 CodeRabbit inference engine (AGENTS.md)
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt,cpp,c,fs}: Write tests before implementation using TDD workflow: write failing test (RED), implement minimal code (GREEN), then refactor (IMPROVE)
Keep functions small (<50 lines) and files focused (<800 lines, typical 200-400 lines)
Avoid deep nesting (>4 levels)
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt}
📄 CodeRabbit inference engine (AGENTS.md)
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt}: Never mutate existing objects; always create new objects with changes applied (Immutability requirement)
Handle errors at every level; provide user-friendly messages in UI code and detailed context in server-side logs
Ensure error messages don't leak sensitive data
Files:
skills/delivery-gate/hooks/quality-gate.py
skills/**/*.{js,ts,py,sh}
📄 CodeRabbit inference engine (AGENTS.md)
Place new workflow contributions in
skills/directory as the canonical workflow surface
Files:
skills/delivery-gate/hooks/quality-gate.py
{skills,commands,agents,rules}/**
⚙️ CodeRabbit configuration file
{skills,commands,agents,rules}/**: Focus on prompt-injection resilience, tool-permission scope, destructive action guards, and secret exfiltration risks.
Files:
skills/delivery-gate/hooks/quality-gate.py
🔇 Additional comments (2)
skills/delivery-gate/hooks/quality-gate.py (2)
58-59: Project-memory key collisions are still possible.The replacement scheme still maps distinct project paths to the same key, so this can select another project’s memory directory. This is the same unresolved concern as the earlier project-memory lookup comment. As per path instructions, “
{skills,commands,agents,rules}/**: Focus on prompt-injection resilience, tool-permission scope, destructive action guards, and secret exfiltration risks.”Source: Path instructions
135-145:transcript_pathstill needs fail-closed validation.
payload['transcript_path']still flows intoopen()without type/root validation, andTypeError/OSErrorare silently swallowed. That can read arbitrary accessible files or downgrade enforcement by falling back to raw stdin. This is the same issue as the earlier transcript-path review. As per coding guidelines, “Never trust external data (API responses, user input, file content)” and “Always handle errors explicitly at every level and never silently swallow errors.” As per path instructions, “{skills,commands,agents,rules}/**: Focus on prompt-injection resilience, tool-permission scope, destructive action guards, and secret exfiltration risks.”Sources: Coding guidelines, Path instructions
| cwd = os.environ.get('CLAUDE_PROJECT_DIR', os.getcwd()) | ||
| safe = cwd.replace(':', '-').replace('\\', '-').replace('/', '-') | ||
| mem = os.path.expanduser(f'~/.claude/projects/{safe}/memory') | ||
| log.info('Looking for memory dir: cwd=%s -> %s', cwd, mem) |
There was a problem hiding this comment.
🔒 Security & Privacy | 🟡 Minor | ⚡ Quick win
Do not log full local project paths at INFO.
Hook stderr is forwarded, and this INFO log can expose usernames, client names, or repo paths in normal hook output. Redact it or demote it to debug.
Proposed fix
- log.info('Looking for memory dir: cwd=%s -> %s', cwd, mem)
+ log.debug('Looking for memory dir for current project')As per path instructions, “{skills,commands,agents,rules}/**: Focus on prompt-injection resilience, tool-permission scope, destructive action guards, and secret exfiltration risks.”
📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| log.info('Looking for memory dir: cwd=%s -> %s', cwd, mem) | |
| log.debug('Looking for memory dir for current project') |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@skills/delivery-gate/hooks/quality-gate.py` at line 60, The INFO log in
quality-gate.py currently prints the full cwd and resolved memory path, which
can leak local project details in normal hook output. Update the logging in the
quality-gate hook to avoid exposing full paths at INFO by either redacting the
path values or moving this message to debug, and keep the change localized to
the hook’s logging call that uses log.info for the memory-dir lookup.
Source: Path instructions
✅ Action performedReview finished.
|
✅ Action performedReviews resumed. |
…uto-closed by fork sync, now on clean branch)
… full transcript (not truncated), memory-dir-absent warning (not silent pass), SKILL.md description accuracy
…plex tasks (fail-close instead of fail-open)
…ipt_path (Greptile feedback)
…de Code actual encoding on Windows
…sh translation to CLAUDE.md block (Greptile feedback)
…new users per Greptile feedback)
…nical patches Reverts 'session hygiene' rebranding. Preserves original approved framing while keeping technical improvements: - JSON transcript_path parsing documentation - filesystem mtime staleness check - 'skip tests for now' rationalization pattern - disk critically low explicit block condition
fff2455 to
ef49724
Compare
✅ Action performedReviews resumed. |
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@skills/delivery-gate/hooks/quality-gate.py`:
- Around line 181-196: The status matrix in quality-gate.py is treating skipped
memory verification as all libraries being fresh because stale is reset to an
empty list when mem_dir is missing. Update the logic around the else branch and
the later parts/status_icons construction so the code preserves an explicit
unchecked state (or omits the matrix entirely) when no memory directory exists,
instead of marking every entry in LIBS as O.
In `@skills/delivery-gate/SKILL.md`:
- Around line 101-119: The blocked examples in SKILL.md use stderr messages that
do not match the actual output from quality-gate.py. Update the examples in the
delivery-gate docs to reflect the current hook messages emitted by the
complex-task and low-disk-space checks, or clearly label them as pseudocode so
readers do not expect strings that cannot appear. Use the existing
quality-gate.py log phrases as the source of truth when revising the example
blocks.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: 3df1d3e3-303f-43fd-8f9a-ab080e00f90a
📒 Files selected for processing (2)
skills/delivery-gate/SKILL.mdskills/delivery-gate/hooks/quality-gate.py
📜 Review details
🧰 Additional context used
📓 Path-based instructions (13)
**/*.{js,ts,jsx,tsx,py,java,cs,go,rb,php,scala,kt}
📄 CodeRabbit inference engine (.cursor/rules/common-coding-style.md)
**/*.{js,ts,jsx,tsx,py,java,cs,go,rb,php,scala,kt}: Always create new objects, never mutate existing ones. Use immutable patterns to prevent hidden side effects and enable safe concurrency
Organize code into many small files (200-400 lines typical, 800 lines max) organized by feature/domain rather than by type
Always handle errors explicitly at every level and never silently swallow errors
Always validate all user input before processing at system boundaries
Use schema-based validation where available
Fail fast with clear error messages when validation fails
Never trust external data (API responses, user input, file content)
Ensure code is readable and well-named
Keep functions small (less than 50 lines)
Keep files focused (less than 800 lines)
Avoid deep nesting (more than 4 levels)
Do not use hardcoded values; use constants or configuration instead
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php,swift,kt,rs,c,cpp,h,hpp}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
No hardcoded secrets (API keys, passwords, tokens) - validate before any commit
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php}: All user inputs must be validated
Enable CSRF protection on all state-changing endpoints
Verify authentication and authorization for all protected endpoints
Implement rate limiting on all endpoints to prevent abuse
Ensure error messages do not leak sensitive data in responses
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php,sql}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
Use parameterized queries to prevent SQL injection
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php,swift,kt,rs,c,cpp,h,hpp,properties,yml,yaml,json,env,config}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
NEVER hardcode secrets in source code - ALWAYS use environment variables or a secret manager
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{py,pyi}
📄 CodeRabbit inference engine (.cursor/rules/python-coding-style.md)
**/*.{py,pyi}: Follow PEP 8 conventions in Python code
Use type annotations on all function signatures in Python
Prefer immutable data structures such as frozen dataclasses and NamedTuple in Python
**/*.{py,pyi}: Auto-format Python files using black/ruff after edit
Run type checking using mypy/pyright after editing Python files
**/*.{py,pyi}: Use Protocol from typing module for duck typing and defining object shapes in Python
Use dataclasses with@dataclassdecorator for DTOs (Data Transfer Objects) in Python
Use context managers (with statement) for resource management in Python
Use generators for lazy evaluation and memory-efficient iteration in Python
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.py
📄 CodeRabbit inference engine (.cursor/rules/python-coding-style.md)
**/*.py: Use black for code formatting in Python
Use isort for import sorting in Python
Use ruff for linting Python codeAvoid using
print()statements in Python code; use theloggingmodule instead
**/*.py: Retrieve secrets and API keys from environment variables using os.environ with error handling (raise KeyError if missing) rather than hardcoding credentials
Use bandit for static security analysis in Python projects
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt,cpp,c,fs}
📄 CodeRabbit inference engine (AGENTS.md)
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt,cpp,c,fs}: Write tests before implementation using TDD workflow: write failing test (RED), implement minimal code (GREEN), then refactor (IMPROVE)
Keep functions small (<50 lines) and files focused (<800 lines, typical 200-400 lines)
Avoid deep nesting (>4 levels)
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt}
📄 CodeRabbit inference engine (AGENTS.md)
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt}: Never mutate existing objects; always create new objects with changes applied (Immutability requirement)
Handle errors at every level; provide user-friendly messages in UI code and detailed context in server-side logs
Ensure error messages don't leak sensitive data
Files:
skills/delivery-gate/hooks/quality-gate.py
skills/**/*.{js,ts,py,sh}
📄 CodeRabbit inference engine (AGENTS.md)
Place new workflow contributions in
skills/directory as the canonical workflow surface
Files:
skills/delivery-gate/hooks/quality-gate.py
{skills,commands,agents,rules}/**
⚙️ CodeRabbit configuration file
{skills,commands,agents,rules}/**: Focus on prompt-injection resilience, tool-permission scope, destructive action guards, and secret exfiltration risks.
Files:
skills/delivery-gate/hooks/quality-gate.pyskills/delivery-gate/SKILL.md
skills/**/*.md
📄 CodeRabbit inference engine (CLAUDE.md)
Skills should be formatted as Markdown with clear sections for When to Use, How It Works, and Examples.
Files:
skills/delivery-gate/SKILL.md
{agents,skills,commands}/**/*.md
📄 CodeRabbit inference engine (CLAUDE.md)
Use lowercase filenames with hyphens (e.g.,
python-reviewer.md,tdd-workflow.md) for agents, skills, and commands.
Files:
skills/delivery-gate/SKILL.md
🧠 Learnings (2)
📚 Learning: 2026-03-15T19:02:43.245Z
Learnt from: imrobinsingh
Repo: affaan-m/everything-claude-code PR: 503
File: skills/data-scraper-agent/SKILL.md:1-748
Timestamp: 2026-03-15T19:02:43.245Z
Learning: In this repository, skill folders should use a lowercase-hyphen name (e.g., data-scraper-agent, claude-api) and the skill description file inside each folder should be named SKILL.md (uppercase). Do not flag SKILL.md as a naming violation; treat SKILL.md as the canonical file name inside each skill directory.
Applied to files:
skills/delivery-gate/SKILL.md
📚 Learning: 2026-04-15T15:52:59.963Z
Learnt from: manja316
Repo: affaan-m/everything-claude-code PR: 1360
File: skills/security-bounty-hunter/SKILL.md:11-18
Timestamp: 2026-04-15T15:52:59.963Z
Learning: In this repository’s skills documentation (skills/**/SKILL.md), use the canonical auto-activation skill section header `## When to Activate`—do not use `## When to Use`. CONTRIBUTING.md and docs/SKILL-DEVELOPMENT-GUIDE.md confirm the required header, and existing skills follow this convention. This header is important for the auto-activation mechanism to detect the correct section.
Applied to files:
skills/delivery-gate/SKILL.md
🪛 ast-grep (0.44.0)
skills/delivery-gate/hooks/quality-gate.py
[warning] 139-139: File path is request-/variable-derived; validate and normalize to prevent path traversal.
Context: open(tp, 'r', encoding='utf-8')
Note: [CWE-22] Improper Limitation of a Pathname to a Restricted Directory ('Path Traversal').
(open-filename-from-request)
[warning] 167-167: Regex pattern passed to re is built from a non-literal (variable, call, concatenation, or f-string) value. If that value is attacker-controlled it can introduce a malicious pattern with catastrophic backtracking (ReDoS). Use a hardcoded literal pattern, or validate/escape untrusted input with re.escape() and bound the regex complexity before compiling.
Context: re.search(p, tail, re.IGNORECASE)
Note: [CWE-1333] Inefficient Regular Expression Complexity.
(redos-non-literal-regex-python)
🪛 Ruff (0.15.18)
skills/delivery-gate/hooks/quality-gate.py
[warning] 74-74: Consider moving this statement to an else block
(TRY300)
[warning] 74-74: Unnecessary assignment to free_gb before return statement
Remove unnecessary assignment
(RET504)
[warning] 80-80: Too many branches (13 > 12)
(PLR0912)
[warning] 85-85: datetime.date.today() used
(DTZ011)
[warning] 96-96: datetime.datetime.fromtimestamp() called without a tz argument
(DTZ006)
[warning] 108-108: datetime.datetime.fromtimestamp() called without a tz argument
(DTZ006)
[warning] 129-129: Too many branches (21 > 12)
(PLR0912)
[warning] 129-129: Too many statements (60 > 50)
(PLR0915)
[warning] 140-140: Unnecessary mode argument
Remove mode argument
(UP015)
[warning] 195-195: zip() without an explicit strict= parameter
Add explicit value for parameter strict=
(B905)
[warning] 207-207: Logging statement uses f-string
(G004)
🔇 Additional comments (4)
skills/delivery-gate/hooks/quality-gate.py (3)
57-60: Use a collision-resistant project-memory key.
cwd.replace(...)is lossy:/work/a-b/cand/work/a/b-cboth resolve to the same~/.claude/projects/...directory. That crosses the privacy boundary here because this hook can read another repo's memory and make allow/block decisions from the wrong project. Use Claude's exact project-dir encoding or another collision-resistant mapping instead. As per path instructions, focus on “tool-permission scope” and “secret exfiltration risks”.Source: Path instructions
135-145: Validatetranscript_pathand fail closed on read errors.This branch trusts
payload['transcript_path'], opens it directly, and silently falls back torawonTypeError/OSError. That both broadens the hook's file-read scope and lets malformed or unreadable paths undercount edits, which can skip the complex-task block. Restrict the path to Claude's transcript root, validate the field type, and treat unsafe or unreadable paths as an explicit failure. As per coding guidelines, “Never trust external data (API responses, user input, file content)” and “Always handle errors explicitly at every level and never silently swallow errors.” As per path instructions, focus on “tool-permission scope” and “secret exfiltration risks”.Sources: Coding guidelines, Path instructions
163-170: Run rationalization checks on the last assistant reply, not the transcript tail.
transcript[-8000:]can match quoted user text or tool output, and it can miss phrases earlier in a long final reply. The Stop payload already carrieslast_assistant_message; use that field forRATIONALIZEinstead of scanning the raw transcript trailer.skills/delivery-gate/SKILL.md (1)
3-3: Document only the checks this hook actually performs.The hook does not detect contradictions, omissions, or unverified assumptions generically, and rationalization hits do not block on their own. These lines currently promise broader “thinking quality” enforcement than
quality-gate.pyimplements.Also applies to: 8-8
| else: | ||
| # No memory dir — setup incomplete. | ||
| # Warn but DO NOT block: blocking here deadlocks new users | ||
| # who haven't created the memory directory yet. | ||
| if is_complex: | ||
| log.warning('No project memory directory found — cannot verify learning capture.') | ||
| log.warning('Set up memory/ per delivery-gate SKILL.md to enable enforcement.') | ||
| stale = [] | ||
|
|
||
| parts = [] | ||
| if is_complex: | ||
| status_icons = ['X' if s in stale else 'O' for s in LIBS] | ||
| parts.append( | ||
| f'\n Complex task ({edit_count} edits). ' | ||
| f'Check: [{"][".join(f"{k}:{v}" for k,v in zip(LIBS.keys(), status_icons))}]' | ||
| ) |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
Don't render all libraries as fresh when memory checks were skipped.
When mem_dir is missing, stale = [], so the status matrix marks every library as O even though Lines 185-187 just said verification was impossible. Keep an explicit unchecked state here, or skip the matrix until a memory directory exists.
Suggested fix
- if is_complex:
+ if is_complex and mem_dir:
status_icons = ['X' if s in stale else 'O' for s in LIBS]
parts.append(
f'\n Complex task ({edit_count} edits). '
f'Check: [{"][".join(f"{k}:{v}" for k,v in zip(LIBS.keys(), status_icons))}]'
)
+ elif is_complex:
+ parts.append(f'\n Complex task ({edit_count} edits). Check: [memory:unavailable]')📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| else: | |
| # No memory dir — setup incomplete. | |
| # Warn but DO NOT block: blocking here deadlocks new users | |
| # who haven't created the memory directory yet. | |
| if is_complex: | |
| log.warning('No project memory directory found — cannot verify learning capture.') | |
| log.warning('Set up memory/ per delivery-gate SKILL.md to enable enforcement.') | |
| stale = [] | |
| parts = [] | |
| if is_complex: | |
| status_icons = ['X' if s in stale else 'O' for s in LIBS] | |
| parts.append( | |
| f'\n Complex task ({edit_count} edits). ' | |
| f'Check: [{"][".join(f"{k}:{v}" for k,v in zip(LIBS.keys(), status_icons))}]' | |
| ) | |
| else: | |
| # No memory dir — setup incomplete. | |
| # Warn but DO NOT block: blocking here deadlocks new users | |
| # who haven't created the memory directory yet. | |
| if is_complex: | |
| log.warning('No project memory directory found — cannot verify learning capture.') | |
| log.warning('Set up memory/ per delivery-gate SKILL.md to enable enforcement.') | |
| stale = [] | |
| parts = [] | |
| if is_complex and mem_dir: | |
| status_icons = ['X' if s in stale else 'O' for s in LIBS] | |
| parts.append( | |
| f'\n Complex task ({edit_count} edits). ' | |
| f'Check: [{"][".join(f"{k}:{v}" for k,v in zip(LIBS.keys(), status_icons))}]' | |
| ) | |
| elif is_complex: | |
| parts.append(f'\n Complex task ({edit_count} edits). Check: [memory:unavailable]') |
🧰 Tools
🪛 Ruff (0.15.18)
[warning] 195-195: zip() without an explicit strict= parameter
Add explicit value for parameter strict=
(B905)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@skills/delivery-gate/hooks/quality-gate.py` around lines 181 - 196, The
status matrix in quality-gate.py is treating skipped memory verification as all
libraries being fresh because stale is reset to an empty list when mem_dir is
missing. Update the logic around the else branch and the later
parts/status_icons construction so the code preserves an explicit unchecked
state (or omits the matrix entirely) when no memory directory exists, instead of
marking every entry in LIBS as O.
…ot stdout
Previously sys.stdout.write(raw) echoed the raw hook JSON payload to stdout,
which Claude Code displays as the hook's response message. When the hook
blocked (exit 2), Claude saw {"transcript_path":"...","session_id":"..."}
instead of the actual blocking reason from stderr.
This made the gate functionally silent from Claude's perspective — it could
not guide Claude to the corrective action (update growth-log / free disk).
Fix per Greptile feedback: stop echo, let stderr messages reach Claude.
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@skills/delivery-gate/hooks/quality-gate.py`:
- Line 176: The warning in quality-gate.py currently logs raw transcript-derived
matches via hits, which can leak sensitive text and re-inject untrusted content
into stderr. Update the logging in the quality-gate feedback path to use a fixed
summary or count only, and avoid including any matched transcript excerpts in
the message. Keep the change localized to the logging statement in the
quality-gate hook logic so the output remains safe and prompt-injection
resistant.
- Around line 153-156: The disk-space blocking warning is emitted twice in the
quality-gate hook, causing duplicate feedback for a single failure. In
quality-gate.py, update the disk space check around the repeated log.warning
call so the message is logged only once, keeping the existing blocked-condition
wording and using the same DISK_CRIT_GB and disk_free context.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: 7469e086-4c33-4ee2-9436-04ae8b5b25f4
📒 Files selected for processing (1)
skills/delivery-gate/hooks/quality-gate.py
📜 Review details
⏰ Context from checks skipped due to timeout. (1)
- GitHub Check: Greptile Review
🧰 Additional context used
📓 Path-based instructions (11)
**/*.{js,ts,jsx,tsx,py,java,cs,go,rb,php,scala,kt}
📄 CodeRabbit inference engine (.cursor/rules/common-coding-style.md)
**/*.{js,ts,jsx,tsx,py,java,cs,go,rb,php,scala,kt}: Always create new objects, never mutate existing ones. Use immutable patterns to prevent hidden side effects and enable safe concurrency
Organize code into many small files (200-400 lines typical, 800 lines max) organized by feature/domain rather than by type
Always handle errors explicitly at every level and never silently swallow errors
Always validate all user input before processing at system boundaries
Use schema-based validation where available
Fail fast with clear error messages when validation fails
Never trust external data (API responses, user input, file content)
Ensure code is readable and well-named
Keep functions small (less than 50 lines)
Keep files focused (less than 800 lines)
Avoid deep nesting (more than 4 levels)
Do not use hardcoded values; use constants or configuration instead
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php,swift,kt,rs,c,cpp,h,hpp}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
No hardcoded secrets (API keys, passwords, tokens) - validate before any commit
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php}: All user inputs must be validated
Enable CSRF protection on all state-changing endpoints
Verify authentication and authorization for all protected endpoints
Implement rate limiting on all endpoints to prevent abuse
Ensure error messages do not leak sensitive data in responses
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php,sql}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
Use parameterized queries to prevent SQL injection
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,cs,rb,go,php,swift,kt,rs,c,cpp,h,hpp,properties,yml,yaml,json,env,config}
📄 CodeRabbit inference engine (.cursor/rules/common-security.md)
NEVER hardcode secrets in source code - ALWAYS use environment variables or a secret manager
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{py,pyi}
📄 CodeRabbit inference engine (.cursor/rules/python-coding-style.md)
**/*.{py,pyi}: Follow PEP 8 conventions in Python code
Use type annotations on all function signatures in Python
Prefer immutable data structures such as frozen dataclasses and NamedTuple in Python
**/*.{py,pyi}: Auto-format Python files using black/ruff after edit
Run type checking using mypy/pyright after editing Python files
**/*.{py,pyi}: Use Protocol from typing module for duck typing and defining object shapes in Python
Use dataclasses with@dataclassdecorator for DTOs (Data Transfer Objects) in Python
Use context managers (with statement) for resource management in Python
Use generators for lazy evaluation and memory-efficient iteration in Python
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.py
📄 CodeRabbit inference engine (.cursor/rules/python-coding-style.md)
**/*.py: Use black for code formatting in Python
Use isort for import sorting in Python
Use ruff for linting Python codeAvoid using
print()statements in Python code; use theloggingmodule instead
**/*.py: Retrieve secrets and API keys from environment variables using os.environ with error handling (raise KeyError if missing) rather than hardcoding credentials
Use bandit for static security analysis in Python projects
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt,cpp,c,fs}
📄 CodeRabbit inference engine (AGENTS.md)
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt,cpp,c,fs}: Write tests before implementation using TDD workflow: write failing test (RED), implement minimal code (GREEN), then refactor (IMPROVE)
Keep functions small (<50 lines) and files focused (<800 lines, typical 200-400 lines)
Avoid deep nesting (>4 levels)
Files:
skills/delivery-gate/hooks/quality-gate.py
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt}
📄 CodeRabbit inference engine (AGENTS.md)
**/*.{js,ts,jsx,tsx,py,java,go,rs,kt}: Never mutate existing objects; always create new objects with changes applied (Immutability requirement)
Handle errors at every level; provide user-friendly messages in UI code and detailed context in server-side logs
Ensure error messages don't leak sensitive data
Files:
skills/delivery-gate/hooks/quality-gate.py
skills/**/*.{js,ts,py,sh}
📄 CodeRabbit inference engine (AGENTS.md)
Place new workflow contributions in
skills/directory as the canonical workflow surface
Files:
skills/delivery-gate/hooks/quality-gate.py
{skills,commands,agents,rules}/**
⚙️ CodeRabbit configuration file
{skills,commands,agents,rules}/**: Focus on prompt-injection resilience, tool-permission scope, destructive action guards, and secret exfiltration risks.
Files:
skills/delivery-gate/hooks/quality-gate.py
🔇 Additional comments (1)
skills/delivery-gate/hooks/quality-gate.py (1)
131-133: LGTM!
| if m: | ||
| hits.append(m.group(0)[:80]) | ||
| if hits: | ||
| log.warning('quality-gate: rationalization detected — %s', hits) |
There was a problem hiding this comment.
🔒 Security & Privacy | 🟡 Minor | ⚡ Quick win
Do not echo transcript excerpts in hook feedback.
hits is transcript-derived. Log a fixed summary/count instead of raw matched text to avoid leaking sensitive content or re-injecting untrusted transcript text via stderr. As per coding guidelines, “Ensure error messages don't leak sensitive data”; as per path instructions, “Focus on prompt-injection resilience, tool-permission scope, destructive action guards, and secret exfiltration risks.”
Proposed fix
if hits:
- log.warning('quality-gate: rationalization detected — %s', hits)
+ log.warning('quality-gate: rationalization detected (%d match/es).', len(hits))📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| log.warning('quality-gate: rationalization detected — %s', hits) | |
| log.warning('quality-gate: rationalization detected (%d match/es).', len(hits)) |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@skills/delivery-gate/hooks/quality-gate.py` at line 176, The warning in
quality-gate.py currently logs raw transcript-derived matches via hits, which
can leak sensitive text and re-inject untrusted content into stderr. Update the
logging in the quality-gate feedback path to use a fixed summary or count only,
and avoid including any matched transcript excerpts in the message. Keep the
change localized to the logging statement in the quality-gate hook logic so the
output remains safe and prompt-injection resistant.
Sources: Coding guidelines, Path instructions
…ot reasoning), warning vs block table, CI/CD analogy, limitations section, self-audit pairing
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@skills/delivery-gate/SKILL.md`:
- Around line 9-126: Add the missing canonical activation section to this skill
doc so the auto-activation mechanism can detect it. In SKILL.md, insert a `##
When to Activate` section in the standard position used by other skills, and
make sure its wording matches the repository convention rather than relying on
`## When to Use` or leaving the section out. Use the existing sections like `##
What It Checks` and `## Why` as anchors when placing it.
- Around line 96-106: The blocked example stderr strings in the delivery gate
docs do not match the actual messages emitted by quality-gate.py. Update the
examples under the complex task and low disk space cases in SKILL.md to use the
exact hook output (or clearly label them as illustrative pseudocode), and verify
the wording against the relevant checks in quality-gate.py such as the
complex-task learning/stale-lib path and the disk-space threshold message.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: 111b501b-092f-4536-9bd0-fd1b7cd7ac91
📒 Files selected for processing (1)
skills/delivery-gate/SKILL.md
📜 Review details
⏰ Context from checks skipped due to timeout. (1)
- GitHub Check: Greptile Review
🧰 Additional context used
📓 Path-based instructions (3)
skills/**/*.md
📄 CodeRabbit inference engine (CLAUDE.md)
Skills should be formatted as Markdown with clear sections for When to Use, How It Works, and Examples.
Files:
skills/delivery-gate/SKILL.md
{agents,skills,commands}/**/*.md
📄 CodeRabbit inference engine (CLAUDE.md)
Use lowercase filenames with hyphens (e.g.,
python-reviewer.md,tdd-workflow.md) for agents, skills, and commands.
Files:
skills/delivery-gate/SKILL.md
{skills,commands,agents,rules}/**
⚙️ CodeRabbit configuration file
{skills,commands,agents,rules}/**: Focus on prompt-injection resilience, tool-permission scope, destructive action guards, and secret exfiltration risks.
Files:
skills/delivery-gate/SKILL.md
🧠 Learnings (2)
📚 Learning: 2026-03-15T19:02:43.245Z
Learnt from: imrobinsingh
Repo: affaan-m/everything-claude-code PR: 503
File: skills/data-scraper-agent/SKILL.md:1-748
Timestamp: 2026-03-15T19:02:43.245Z
Learning: In this repository, skill folders should use a lowercase-hyphen name (e.g., data-scraper-agent, claude-api) and the skill description file inside each folder should be named SKILL.md (uppercase). Do not flag SKILL.md as a naming violation; treat SKILL.md as the canonical file name inside each skill directory.
Applied to files:
skills/delivery-gate/SKILL.md
📚 Learning: 2026-04-15T15:52:59.963Z
Learnt from: manja316
Repo: affaan-m/everything-claude-code PR: 1360
File: skills/security-bounty-hunter/SKILL.md:11-18
Timestamp: 2026-04-15T15:52:59.963Z
Learning: In this repository’s skills documentation (skills/**/SKILL.md), use the canonical auto-activation skill section header `## When to Activate`—do not use `## When to Use`. CONTRIBUTING.md and docs/SKILL-DEVELOPMENT-GUIDE.md confirm the required header, and existing skills follow this convention. This header is important for the auto-activation mechanism to detect the correct section.
Applied to files:
skills/delivery-gate/SKILL.md
🔇 Additional comments (2)
skills/delivery-gate/SKILL.md (2)
3-4: LGTM!
23-23: Verify: Does rationalization regex still scan "transcript tail"?The docs say "Regex on transcript tail" but
count_editswas changed to scan the full transcript per PR fixes. Confirm whether rationalization detection also uses the full transcript or still reads only the tail. If the latter, the distinction is worth noting; if the former, update to "full transcript."
| # Delivery Gate — Mechanical Quality Gate for Claude Code | ||
|
|
||
| A **Stop hook** that checks three things before Claude can finish a session, using only **deterministic checks** — file modification timestamps, disk usage, and regex patterns on the transcript text. No AI inference. | ||
|
|
||
| This is distinct from reasoning gates (like `self-audit`): delivery-gate checks machine-verifiable facts; self-audit checks output quality across four reasoning dimensions. Together they form defense in depth: | ||
| - **delivery-gate**: "Was the learning library touched today? Is disk space safe?" | ||
| - **self-audit**: "Is the file content correct, complete, and honest?" | ||
|
|
||
| This is the same pattern as CI pipeline gates — automated, deterministic checks that verify machine-readable facts rather than trusting self-reported status. | ||
|
|
||
| ## What It Checks | ||
|
|
||
| | Check | Mechanism | On Hit | | ||
| |-------|-----------|--------| | ||
| | Rationalization patterns | Regex on transcript tail | **Warning only** (never blocks) | | ||
| | Stale learning libraries | mtime on 5 configurable paths | Warning if some stale; **Block** if ALL stale + complex task | | ||
| | Disk space < 50GB | `shutil.disk_usage` | Warning | | ||
| | Disk space < 15GB | `shutil.disk_usage` | **Block** (exit 2) | | ||
|
|
||
| Rationalization detection warns about patterns like "skip tests for now" and "pre-existing bug" — surface signals that thinking may have been cut short. It never blocks on its own, because regex heuristics can false-positive. The blocking conditions are only: disk critical OR all learning libraries untouched after a complex task. | ||
|
|
||
| ## Why | ||
|
|
||
| Claude Code's built-in checks cover code quality (build → type → lint → test). But there's a different failure mode: the agent produces working code while the **session hygiene was neglected** — learning not captured, rationalized shortcuts, disk running out silently. | ||
|
|
||
| Over many sessions of "ship and forget," the human hasn't grown. This hook enforces the habit: complex task → must touch learning libraries. | ||
|
|
||
| ## Install | ||
|
|
||
| ```bash | ||
| cp quality-gate.py ~/.claude/scripts/ | ||
| ``` | ||
|
|
||
| Add to `~/.claude/settings.json`: | ||
| ```json | ||
| { | ||
| "hooks": { | ||
| "Stop": [{ | ||
| "hooks": [{ | ||
| "type": "command", | ||
| "command": "python3 ~/.claude/scripts/quality-gate.py", | ||
| "timeout": 5000 | ||
| }] | ||
| }] | ||
| } | ||
| } | ||
| ``` | ||
|
|
||
| ## Learning Libraries | ||
|
|
||
| Create these files in your project's memory directory. The hook checks if at least one was updated today: | ||
|
|
||
| ``` | ||
| memory/ | ||
| ├── growth-log/ # Daily learning entries (directory) | ||
| ├── decisions/log.md # Decision log | ||
| ├── output-index.md # Index of session outputs | ||
| ├── ratings-tracker.md # Skill ratings over time | ||
| └── tooling_capabilities.md # Known tools inventory | ||
| ``` | ||
|
|
||
| Customize the `LIBS` dict to match your own file structure. | ||
|
|
||
| ## Configuration | ||
|
|
||
| Edit `quality-gate.py`: | ||
|
|
||
| | Variable | Default | Purpose | | ||
| |----------|---------|---------| | ||
| | `RATIONALIZE` | 4 patterns | Regex patterns for rationalization detection | | ||
| | `LIBS` | 5 libraries | Files/dirs to check for today's updates | | ||
| | `COMPLEX_THRESHOLD` | 3 | Edit/Write calls to classify as complex | | ||
| | `DISK_WARN_GB` | 50 | Warn below this | | ||
| | `DISK_CRIT_GB` | 15 | Block below this | | ||
|
|
||
| ## Examples | ||
|
|
||
| **Simple session — allowed:** | ||
| ``` | ||
| edit_count=1 (< 3, not complex) → exit 0 | ||
| ``` | ||
|
|
||
| **Complex task, learning captured — allowed:** | ||
| ``` | ||
| edit_count=5 (complex) → checks LIBS → growth-log updated today → exit 0 | ||
| ``` | ||
|
|
||
| **Complex task, no learning — BLOCKED:** | ||
| ``` | ||
| edit_count=4 (complex) → checks LIBS → all 5 stale → exit 2 | ||
| stderr: "Blocked: complex task completed but no learning captured today." | ||
| ``` | ||
|
|
||
| **Low disk space — BLOCKED:** | ||
| ``` | ||
| disk_free=12GB < 15GB critical → exit 2 | ||
| stderr: "Blocked: disk space at 12GB (threshold: 15GB)." | ||
| ``` | ||
|
|
||
| ## Limitations | ||
|
|
||
| The hook enforces the **habit** of touching learning libraries, not the **quality** of what was recorded. If `output-index.md` is updated but `growth-log` is skipped, the hook passes (1 of 5 libraries touched). This is by design: mechanical gates check machine-verifiable facts. For content quality verification, pair with `self-audit`. | ||
|
|
||
| ## Compatibility | ||
|
|
||
| - Python 3.8+ (uses `from __future__ import annotations`) | ||
| - Cross-platform: Windows, macOS, Linux | ||
| - Zero dependencies beyond stdlib | ||
|
|
||
| ## Quality | ||
|
|
||
| This code went through 4 rounds of automated code review (CodeRabbit + Greptile) with 9 real bugs found and fixed. | ||
|
|
||
| ## See Also | ||
|
|
||
| - `self-audit` — Reasoning quality gate (completeness/consistency/groundedness/honesty) | ||
| - `verification-loop` — Code quality checks (build/type/lint/test) | ||
| - `gateguard` — PreToolUse safety gate |
There was a problem hiding this comment.
📐 Maintainability & Code Quality | 🟠 Major | ⚡ Quick win
Missing required ## When to Activate section.
Per this repository's skill documentation convention, the canonical auto-activation header is ## When to Activate, not ## When to Use or omission. This section is required for the auto-activation mechanism to detect the correct section. Based on learnings, CONTRIBUTING.md and docs/SKILL-DEVELOPMENT-GUIDE.md confirm this requirement, and existing skills follow this convention.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@skills/delivery-gate/SKILL.md` around lines 9 - 126, Add the missing
canonical activation section to this skill doc so the auto-activation mechanism
can detect it. In SKILL.md, insert a `## When to Activate` section in the
standard position used by other skills, and make sure its wording matches the
repository convention rather than relying on `## When to Use` or leaving the
section out. Use the existing sections like `## What It Checks` and `## Why` as
anchors when placing it.
Source: Learnings
| **Complex task, no learning — BLOCKED:** | ||
| ``` | ||
| edit_count=4 (complex) → checks LIBS → all 5 stale → exit 2 | ||
| stderr: "Blocked: complex task completed but no learning captured today." | ||
| ``` | ||
|
|
||
| **Low disk space — BLOCKED:** | ||
| ``` | ||
| disk_free=12GB < 15GB critical → exit 2 | ||
| stderr: "Blocked: disk space at 12GB (threshold: 15GB)." | ||
| ``` |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
Example stderr messages still don't match actual hook output.
The blocked examples quote strings quality-gate.py never emits:
- Docs:
"Blocked: complex task completed but no learning captured today." - Actual:
"Blocked: complex task but >=3 learning libs stale."or"Blocked: code changes made but no growth-log update." - Docs:
"Blocked: disk space at 12GB (threshold: 15GB)." - Actual:
"Blocked: disk space at %dGB (<%dGB). Free space before continuing."
Update these to match quality-gate.py log output or mark as illustrative pseudocode.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@skills/delivery-gate/SKILL.md` around lines 96 - 106, The blocked example
stderr strings in the delivery gate docs do not match the actual messages
emitted by quality-gate.py. Update the examples under the complex task and low
disk space cases in SKILL.md to use the exact hook output (or clearly label them
as illustrative pseudocode), and verify the wording against the relevant checks
in quality-gate.py such as the complex-task learning/stale-lib path and the
disk-space threshold message.
…atch "we can fix" and "integration tests" variants
v1.1.1 — documentation refined + regex expanded + fully testedWhat changed since daltino's approval on #2365Documentation (SKILL.md):
Code (quality-gate.py):
Test Results — 4 categories, 27 scenarios25/27 pass. 2 known regex coverage limits (casual phrasing), not bugs. What was learned (testing CL hooks)The hook counts edits from JSON tool-call patterns in the transcript ( Relationship to #2365Same contribution. daltino's approval applies. v1.1.1 adds regex coverage + doc accuracy, no functional changes to the approved hook logic. |
There was a problem hiding this comment.
Actionable comments posted: 1
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (2)
skills/delivery-gate/SKILL.md (2)
1-9: 🎯 Functional Correctness | 🔴 Critical | ⚡ Quick winMissing required
## When to Activatesection.The canonical auto-activation header
## When to Activateis absent. Per this repository's convention, this section is required for the auto-activation mechanism to detect the skill. The file jumps straight from metadata to# Delivery Gatewithout the standard activation section. Based on learnings, CONTRIBUTING.md and docs/SKILL-DEVELOPMENT-GUIDE.md confirm this requirement, and existing skills follow this convention.Add
## When to Activatein the standard position, typically after the description and before## What It Checksor## Why.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@skills/delivery-gate/SKILL.md` around lines 1 - 9, The delivery-gate skill is missing the required auto-activation header, so update the SKILL.md content to include a standard “## When to Activate” section in the expected spot after the metadata/intro and before the next sections. Use the existing delivery-gate headings like “# Delivery Gate” and any following section headers to place it consistently with other skills, matching the repository’s activation convention.Source: Learnings
57-71: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win"At least one was updated today" misrepresents the check.
The hook evaluates each library individually for today's mtime; it does not pass on "at least one" updated. A session with 1 of 5 libraries updated still has 4 stale, which can block if the task is complex and growth-log is among the stale. Replace with language that matches the per-library staleness evaluation, e.g., "The hook checks which of these libraries were updated today."
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@skills/delivery-gate/SKILL.md` around lines 57 - 71, The “Learning Libraries” text misstates the hook’s freshness check by saying “at least one was updated today,” which implies a pass/fail based on any single file; update the wording in SKILL.md to reflect that the hook evaluates each library’s mtime individually and identifies which specific libraries were updated today. Keep the guidance aligned with the existing LIBS list and the memory/ directory examples so the description matches the actual per-library staleness behavior.
♻️ Duplicate comments (1)
skills/delivery-gate/SKILL.md (1)
96-106: 🎯 Functional Correctness | 🔴 Critical | ⚡ Quick winExample stderr messages still do not match actual hook output.
These fabricated strings persist from prior review rounds:
Line 99 claims:
"Blocked: complex task completed but no learning captured today."
- Actual:
"Blocked: complex task but >=3 learning libs stale."Line 105 claims:
"Blocked: disk space at 12GB (threshold: 15GB)."
- Actual:
"Blocked: disk space at %dGB (<%dGB). Free space before continuing."Users debugging against these examples will search for strings that never appear. Either update to the exact
quality-gate.pylog output or clearly label these as illustrative pseudocode. This was previously flagged and marked addressed but remains unfixed.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@skills/delivery-gate/SKILL.md` around lines 96 - 106, The example stderr strings in SKILL.md do not match the real hook output, so update the examples to mirror the exact messages emitted by quality-gate.py, or clearly mark them as pseudocode; specifically, revise the BLOCKED examples under the complex task and low disk space cases so they use the actual stderr text from the hook output rather than fabricated strings.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@skills/delivery-gate/SKILL.md`:
- Around line 21-27: The delivery-gate severity table is inconsistent with the
actual behavior in the disk-space check. Update the SKILL.md guidance for the
disk thresholds so the entry aligned with the 50GB reminder matches the
implementation in the delivery gate logic, and reflect the distinct
INFO/WARNING/BLOCK levels used by the disk check. Locate the disk-space section
by the "Disk space < 50GB" and "Disk space < 15GB" rows and revise the "On Hit"
labels so they accurately describe the reminder behavior instead of calling it a
warning.
---
Outside diff comments:
In `@skills/delivery-gate/SKILL.md`:
- Around line 1-9: The delivery-gate skill is missing the required
auto-activation header, so update the SKILL.md content to include a standard “##
When to Activate” section in the expected spot after the metadata/intro and
before the next sections. Use the existing delivery-gate headings like “#
Delivery Gate” and any following section headers to place it consistently with
other skills, matching the repository’s activation convention.
- Around line 57-71: The “Learning Libraries” text misstates the hook’s
freshness check by saying “at least one was updated today,” which implies a
pass/fail based on any single file; update the wording in SKILL.md to reflect
that the hook evaluates each library’s mtime individually and identifies which
specific libraries were updated today. Keep the guidance aligned with the
existing LIBS list and the memory/ directory examples so the description matches
the actual per-library staleness behavior.
---
Duplicate comments:
In `@skills/delivery-gate/SKILL.md`:
- Around line 96-106: The example stderr strings in SKILL.md do not match the
real hook output, so update the examples to mirror the exact messages emitted by
quality-gate.py, or clearly mark them as pseudocode; specifically, revise the
BLOCKED examples under the complex task and low disk space cases so they use the
actual stderr text from the hook output rather than fabricated strings.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: 714e8c5f-5ad6-4d0c-9794-3f5b2788fe08
📒 Files selected for processing (1)
skills/delivery-gate/SKILL.md
📜 Review details
⏰ Context from checks skipped due to timeout. (1)
- GitHub Check: Greptile Review
🧰 Additional context used
📓 Path-based instructions (3)
skills/**/*.md
📄 CodeRabbit inference engine (CLAUDE.md)
Skills should be formatted as Markdown with clear sections for When to Use, How It Works, and Examples.
Files:
skills/delivery-gate/SKILL.md
{agents,skills,commands}/**/*.md
📄 CodeRabbit inference engine (CLAUDE.md)
Use lowercase filenames with hyphens (e.g.,
python-reviewer.md,tdd-workflow.md) for agents, skills, and commands.
Files:
skills/delivery-gate/SKILL.md
{skills,commands,agents,rules}/**
⚙️ CodeRabbit configuration file
{skills,commands,agents,rules}/**: Focus on prompt-injection resilience, tool-permission scope, destructive action guards, and secret exfiltration risks.
Files:
skills/delivery-gate/SKILL.md
🧠 Learnings (2)
📚 Learning: 2026-03-15T19:02:43.245Z
Learnt from: imrobinsingh
Repo: affaan-m/everything-claude-code PR: 503
File: skills/data-scraper-agent/SKILL.md:1-748
Timestamp: 2026-03-15T19:02:43.245Z
Learning: In this repository, skill folders should use a lowercase-hyphen name (e.g., data-scraper-agent, claude-api) and the skill description file inside each folder should be named SKILL.md (uppercase). Do not flag SKILL.md as a naming violation; treat SKILL.md as the canonical file name inside each skill directory.
Applied to files:
skills/delivery-gate/SKILL.md
📚 Learning: 2026-04-15T15:52:59.963Z
Learnt from: manja316
Repo: affaan-m/everything-claude-code PR: 1360
File: skills/security-bounty-hunter/SKILL.md:11-18
Timestamp: 2026-04-15T15:52:59.963Z
Learning: In this repository’s skills documentation (skills/**/SKILL.md), use the canonical auto-activation skill section header `## When to Activate`—do not use `## When to Use`. CONTRIBUTING.md and docs/SKILL-DEVELOPMENT-GUIDE.md confirm the required header, and existing skills follow this convention. This header is important for the auto-activation mechanism to detect the correct section.
Applied to files:
skills/delivery-gate/SKILL.md
🔇 Additional comments (1)
skills/delivery-gate/SKILL.md (1)
108-126: LGTM!
| | Check | Mechanism | On Hit | | ||
| |-------|-----------|--------| | ||
| | Rationalization patterns | Regex on transcript tail | **Warning only** (never blocks) | | ||
| | Stale learning libraries | mtime on 5 configurable paths | Warning if some stale; **Block** if >=3 stale OR growth-log stale + complex task | | ||
| | Disk space < 50GB | `shutil.disk_usage` | Warning | | ||
| | Disk space < 15GB | `shutil.disk_usage` | **Block** (exit 2) | | ||
|
|
There was a problem hiding this comment.
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
Disk <50GB is labeled "Warning" but code emits INFO level reminder.
The implementation uses three severity levels: <50GB logs INFO ("Reminder"), <30GB logs WARNING, and <15GB blocks. The table conflates the 50GB remind threshold with a warning, which misrepresents the actual user experience. Update the "On Hit" column for the 50GB row to "Reminder (INFO)" or restructure to show all three levels accurately.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@skills/delivery-gate/SKILL.md` around lines 21 - 27, The delivery-gate
severity table is inconsistent with the actual behavior in the disk-space check.
Update the SKILL.md guidance for the disk thresholds so the entry aligned with
the 50GB reminder matches the implementation in the delivery gate logic, and
reflect the distinct INFO/WARNING/BLOCK levels used by the disk check. Locate
the disk-space section by the "Disk space < 50GB" and "Disk space < 15GB" rows
and revise the "On Hit" labels so they accurately describe the reminder behavior
instead of calling it a warning.
| - Zero dependencies beyond stdlib | ||
|
|
||
| ## Quality | ||
|
|
There was a problem hiding this comment.
Limitations example contradicts the blocking logic
The prose says "If output-index.md is updated but growth-log is skipped, the hook passes (1 of 5 libraries touched)." This is wrong. The code at lines 200–206 of quality-gate.py has a dedicated check: if 'growth-log' in stale: sys.exit(2), independent of how many other libraries were touched. A complex session that updates every library except growth-log will be blocked. A user who reads this section and acts accordingly (updating only output-index.md) will find themselves unexpectedly stuck. The example should read: "If growth-log is updated but the other 4 libraries are skipped, the hook passes."
* chore(catalog): sync manifests after skill batch (#2275 #2377 #2378 #2381) Update skill counts (273 -> 277) across catalog docs after the verified skill batch. * fix(skills): replace emoji with ASCII in growth-log + loop-design-check check-unicode-safety (pre-push gate) bans emoji in SKILL.md; the merged #2377 and #2381 slipped through run-all.js. Swap U+274C/U+2705 for 'Avoid:'/'Bad:'/'Good:'.
Install delivery-gate as a Stop hook (fail-open) so users get learning-capture enforcement on session end without manual configuration. Expose delivery-gate and growth-log skills through install-modules and package.json publish surface. Builds on merged affaan-m#2377 (growth-log methodology) and affaan-m#2378 (delivery-gate Python hook), porting to Node.js and closing the last mile: zero-config auto-activation. 6 files, +643/-2
Install delivery-gate as a Stop hook (fail-open) so users get learning-capture enforcement on session end without manual configuration. Expose delivery-gate and growth-log skills through install-modules and package.json publish surface. Builds on merged affaan-m#2377 (growth-log methodology) and affaan-m#2378 (delivery-gate Python hook), porting to Node.js and closing the last mile: zero-config auto-activation. 6 files, +643/-2
…, assumptions, stale logs, disk space (delivery-gate) (affaan-m#2378) * Restore delivery-gate: Stop hook with learning capture enforcement (auto-closed by fork sync, now on clean branch) * Fix bot findings: log level→INFO (DISK_REMIND dead code), count_edits full transcript (not truncated), memory-dir-absent warning (not silent pass), SKILL.md description accuracy * Fix CodeRabbit feedback: treat missing memory-dir as all-stale on complex tasks (fail-close instead of fail-open) * Trigger bot re-review (no logic changes) * Fix: handle both stdin formats — raw transcript AND JSON with transcript_path (Greptile feedback) * Add debug log for memory-dir lookup path * Fix path encoding: replace colon with dash (not strip), matching Claude Code actual encoding on Windows * Fix SKILL.md: update How It Works for JSON+transcript_path, add English translation to CLAUDE.md block (Greptile feedback) * Fix: memory-dir absent → warn but don't block (prevents deadlock for new users per Greptile feedback) * fix: restore daltino-approved voice (thinking quality/收尾铁律) with technical patches Reverts 'session hygiene' rebranding. Preserves original approved framing while keeping technical improvements: - JSON transcript_path parsing documentation - filesystem mtime staleness check - 'skip tests for now' rationalization pattern - disk critically low explicit block condition * fix: remove stdout JSON echo — Stop hooks write feedback to stderr, not stdout Previously sys.stdout.write(raw) echoed the raw hook JSON payload to stdout, which Claude Code displays as the hook's response message. When the hook blocked (exit 2), Claude saw {"transcript_path":"...","session_id":"..."} instead of the actual blocking reason from stderr. This made the gate functionally silent from Claude's perspective — it could not guide Claude to the corrective action (update growth-log / free disk). Fix per Greptile feedback: stop echo, let stderr messages reach Claude. * fix: remove duplicate disk-critical log line * docs(delivery-gate): v1.1.0 — accurate scope (deterministic checks, not reasoning), warning vs block table, CI/CD analogy, limitations section, self-audit pairing * fix(delivery-gate): expand rationalization regex coverage (R3/R4) — match "we can fix" and "integration tests" variants * chore: bump version to 1.1.1 to re-trigger CI checks
Install delivery-gate as a Stop hook (fail-open) so users get learning-capture enforcement on session end without manual configuration. Expose delivery-gate and growth-log skills through install-modules and package.json publish surface. Builds on merged affaan-m#2377 (growth-log methodology) and affaan-m#2378 (delivery-gate Python hook), porting to Node.js and closing the last mile: zero-config auto-activation. 6 files, +643/-2
Install delivery-gate as a Stop hook (fail-open) so users get learning-capture enforcement on session end without manual configuration. Expose delivery-gate and growth-log skills through install-modules and package.json publish surface. Builds on merged affaan-m#2377 (growth-log methodology) and affaan-m#2378 (delivery-gate Python hook), porting to Node.js and closing the last mile: zero-config auto-activation. 6 files, +643/-2
* docs: add MRR-biased ECC Pro + AgentShield security roadmap Output of a multi-agent survey + research pass: capability map of AgentShield and ECC Pro, triage of every open PR/issue on both repos, and web research on competitors, unbuilt ideas, and dev-tool demand. 17 items across 4 themes (now/next/later) scored for free-to-paid conversion, each linked to the real PRs/issues that implement it. Includes the reusable workflow script that generated it. Headline: ecc-agentshield is ~30K downloads/month with near-zero monetization bridge, and the agent-proximity moat is computed but never rendered. Roadmap removes trust blockers (FP cluster), makes the moat visible (PR #2320), then productizes local CLI primitives into hosted Pro surfaces. * docs(design): add hosted Pro fleet dashboard design (Sentry for agent security) Implementation-ready architecture for the flagship 'next' roadmap item: a hosted, multi-repo agent-security posture dashboard built on the existing ecc-agentshield primitives (evidence-pack bundleDigest + operatorReadback, watch/drift DriftResult, runtime NDJSON, baseline diff, policy promotion). Covers free-vs-Pro scope, ingestion/query API grounded in real field names, data model + time-series rollups, auth/RBAC + redaction guarantees, MVP build order, and pricing hooks. Companion to ECC-PRO-SECURITY-ROADMAP.md. * feat(control-pane): serve 3D agent-airspace viz + /api/proximity feed (#2320) Adds the Layer 4 observability view to the control pane: a self-contained, dependency-free 3D point-cloud of the agent airspace (positions from the proximity embedding, sized by working set, colored by collision risk, links for converging pairs) plus an XSS-safe advisory panel that polls every 5s. - proximity-viz.js: renderProximityVizHtml() (canvas projection, no external JS) - server.js: GET /proximity (page) + GET /api/proximity (snapshot.proximity feed) - test: asserts both routes serve and the feed carries positions/links/advisories * fix(clv2): escape $HOME before pgrep -f in migrate-homunculus.sh (#2339) * fix(clv2): escape $HOME before pgrep -f in migrate-homunculus.sh pgrep -f treats its argument as an extended regular expression, but the running-observer guard interpolated $HOME unescaped. Paths containing regex metacharacters (e.g. /home/user.name, /home/c++dev, /home/user (work)) made the match over-broad or invalid, causing either a false negative (live observer missed, migration proceeds and risks registry corruption) or a false positive (migration blocked unnecessarily). Escape the ERE metacharacters in $HOME via sed before building the pattern so the home prefix is matched literally while the trailing .*observer-loop\.sh regex is preserved. Portable across BSD and GNU sed. Fixes #2301 * test(clv2): add regression test for migrate-homunculus.sh $HOME escaping Guards the #2301 fix: extracts the script's sed escaping command and asserts the resulting pgrep -f pattern matches the literal home path while no longer over-matching a regex-expanded decoy (HOME=/home/user.name must not match /home/userXname). Also pins that the guard uses escaped_home rather than $HOME directly. Follows the existing clv2 shell-test convention in tests/hooks/observe-entrypoint-allowlist.test.js. Refs #2301 * test(clv2): skip migrate-homunculus escaping test on Windows The test relies on POSIX bash/sed/grep -E semantics, which differ on the Windows CI runners. Guard with the same process.platform === 'win32' early exit used by tests/hooks/observe-subdirectory-detection.test.js so the bash-dependent assertions only run on POSIX platforms. Refs #2301 * fix(clv2): harden registry writes and project deletion (#2294, #2297) (#2323) Two security-priority fixes in continuous-learning-v2/scripts/instinct-cli.py: - #2294: _write_registry wrote projects.json without the advisory lock that _update_registry holds, so concurrent 'projects delete/gc/merge' could race an observe-time update and corrupt the registry. Extract the lock into a shared _registry_lock() context manager and use it in both writers. - #2297: _remove_project_storage called shutil.rmtree on PROJECTS_DIR/project_id with no containment check. Add defense-in-depth: resolve the path and refuse to delete anything that is not strictly inside PROJECTS_DIR (or is the root itself), so a relaxed validator or future caller can never cause an arbitrary-directory delete. Adds 5 pytest regression tests (atomic write under lock, contained delete, missing-dir no-op, traversal refused, root refused). Node integration suite (tests/scripts/instinct-cli-projects.test.js) green 9/9. * feat(workflows): add orch-review native Workflow pilot (#2363) * feat(workflows): add orch-review native Workflow pilot Port orch-pipeline Phase 5 (Review) to a native Claude Code Workflow script. The gated outer loop stays in the main conversation; this script owns only the autonomous review+verify segment between the two human gates: 1. Review — reviewers fan out in parallel: ecc:code-reviewer always, ecc:<language>-reviewer when args.language maps, ecc:security-reviewer when the orch-pipeline security trigger matches the diff/paths. 2. Dedup — merge findings across dimensions keyed on the normalized evidence snippet, since independent reviewers flag the same line. 3. Verify — each unique CRITICAL/HIGH finding goes to an independent adversarial verifier; MEDIUM/LOW pass through as advisory. The Review->Verify barrier is deliberate: deduping before verification stops the verifier running N times on the same bug (local testing: 11 raw findings collapsed to 4 unique, ~halving verifier cost). Existing ECC reviewer subagents are reused via agentType; reviewer output is validated by JSON schema. args is accepted as an object or a JSON-encoded string. - workflows/orch-review.workflow.js — the workflow script - workflows/README.md — invocation contract, returns shape, follow-ups CI lint is scoped to scripts/ and tests/, so the script (validated with node --check) and the README (passes markdownlint) are untouched. * fix(workflows): fail closed on invalid args and lost review dimensions Addresses the two safety findings from the PR bot review: 1. Lost review dimension (Greptile P1 / CodeRabbit Major): a reviewer agent that returns null or rejects was silently dropped by filter(Boolean), so an unreviewed security dimension could still return APPROVE. Each dimension's outcome is now captured; failures land in failedDimensions and force CHANGES_REQUESTED (incomplete). 2. Invalid args (CodeRabbit Major): an empty diff returned APPROVE and bad JSON / non-array changedFiles threw inconsistently. Input is now validated up front and rejected with a clear error — the gate fails closed instead of approving an unreviewed payload. Docs (header contract + README) updated for the new return fields (incomplete, failedDimensions, stats.failed). Remaining bot nits (evidence minLength, verify-label collision, verified->confirmed rename, contract drift) deferred as follow-ups. * fix(workflows): address remaining orch-review review nits Follow-up to the bot review (deferred items from the safety pass): - evidence: require minLength 1 in the schema, and fall back to a title+line dedup key when evidence is empty, so empty-evidence findings in one file no longer collapse onto a single key and drop (CodeRabbit). - verify label: include a slice of the normalized evidence so two CRITICAL/HIGH findings from the same file get distinct labels and do not alias under resumability (Greptile). - stats.verified -> stats.confirmed to match the "confirmed" wording used in the log and avoid ambiguity vs the refuted count (Greptile); header contract and README updated to match. Verified by running the workflow on a synthetic vulnerable diff: dedup 12 raw -> 5 unique, stats.confirmed populated, fail-closed fields (incomplete/failedDimensions) intact. * fix(workflows): harden verify stage and diff-only verification Addresses the second-round bot review: - Verify stage now has the same failure guard as the review stage: a rejected verifier no longer nulls out its slot (which crashed the later filter). A null return is treated as unconfirmed; a rejection keeps the finding as blocking (fail closed) so an unverifiable CRITICAL is never silently demoted to advisory (CodeRabbit @221). - verifyPrompt now instructs the skeptic to judge solely from the provided diff text and not to refute merely because the referenced file is absent from the working tree (the diff may be an unapplied PR). Fixes the false-refute seen when testing on a synthetic diff. CodeRabbit @81 (evidence minLength) was already addressed in the prior commit; this is a stale re-post on the unresolved thread. * fix(workflows): keep unverifiable blockers blocking; stop leaking error text Second-round bot review (CodeRabbit): - @218 Treat a null/failed verifier as `unverified`, not refuted. A terminal verifier failure or skip no longer demotes a CRITICAL/HIGH to advisory; it stays in `blocking` tagged "could not be verified" (fail closed). Only a genuine isReal=false verdict is refuted. Adds stats.unverified. - @189 Do not return raw subagent error text. Review/verify failures now log the raw message for operators and return only a bounded label (failedDimensions[].error = "review agent failed"). Stale re-posts this round (@81 evidence minLength, @224 verify guard) were already fixed in prior commits. * docs(workflows): enumerate bounded failedDimensions.error labels CodeRabbit (trivial): the public contract implied callers get human-readable error text, but the implementation returns only bounded labels. Enumerate them in the README returns block. * Update yarn.lock (#2342) * Add memxus configuration to mcp-servers.json (#2355) * Add memxus configuration to mcp-servers.json Added configuration for Memxus service with API key placeholder and description. * Revise description in mcp-servers.json Updated the description to include a note about reviewing stored memories to prevent prompt-injection. * Update description in mcp-servers.json Update description in mcp-servers.json * Fix for docs: Scope Decision Guide table duplicated in SKILL.md and observer.md with minor drift (#2366) #2306 Co-authored-by: angadsingh7666 <[email protected]> * fix(llm): align Claude provider with current Anthropic API (#2133) Replace invalid default model IDs (e.g. claude-sonnet-4-7) with current claude-sonnet-4-6, claude-opus-4-8, and claude-haiku-4-5. Route system messages to the API system field, enable ephemeral prompt caching, omit temperature for Opus 4.7/4.8, and surface cache usage metrics. Update the CLI model picker to match. Co-authored-by: Vladimir Đuranović <[email protected]> Co-authored-by: Cursor <[email protected]> * fix(release): derive approval gate paths from version (#2383) Co-authored-by: jan <[email protected]> * fix(release): derive video suite paths from version (#2384) Co-authored-by: jan <[email protected]> * ci: isolate OMP workflow verification (#2382) Co-authored-by: jan <[email protected]> * fix(tests): resolve 10 failing tests on Windows (#2307) - resolve-formatter: stop findProjectRoot walk before os.homedir() to avoid mistaking global dotfiles (e.g. ~/.prettierrc) for a project root - instinct-cli-projects: detect python3/python binary at runtime; skip gracefully when Python 3 is unavailable instead of crashing with null status - command-registry: regenerate COMMAND-REGISTRY.json (was stale) Co-authored-by: Claude Sonnet 4.6 <[email protected]> * fix(hooks): quote args when probing Windows .cmd MCP servers via shell (#2343) On Windows, when a bare-name MCP server command (e.g. codesys-mcp-sp21-plus) falls back to the .cmd candidate, the probe sets shell:true to work around Node 18.20+ CVE-2024-27980. However, passing an args array alongside shell:true causes Node to concatenate the tokens without quoting (DEP0190), so an arg containing a space (e.g. --codesys-path "C:\Program Files\...") is re-split by cmd.exe at every space boundary. The child process receives a truncated path, fails to launch, and the probe declares the server unavailable, falsely blocking every MCP tool call to that server. Fix: add a quoteWin() helper that double-quotes any token containing whitespace or cmd metacharacters. In the useShell branch, build a single properly-quoted command line string and pass it as the sole argument to spawn() with no separate args array. The else branch (shell:false, all non-.cmd commands) is unchanged. Regression test added: on Windows, creates a .cmd shim that echoes its first positional argument to stderr, probes it with a space-containing path arg, and asserts the probe succeeds and the arg was not split at the space boundary. Co-authored-by: Karstein Phobic Nyvold Kvistad <[email protected]> * fix(hooks): guard doc-file-warning stdin listeners behind require.main (#2358) * fix(hooks): guard doc-file-warning stdin listeners behind require.main doc-file-warning.js registered process.stdin data/end listeners at module scope while also exporting run(). run-with-flags.js require()s any hook that exports run() for its in-process fast path, so importing this hook attached stray stdin listeners to the dispatcher process, corrupting the PreToolUse stdout JSON contract. This is the exact failure run-with-flags' own SAFETY comment warns about, and 24 sibling hooks already guard against it. - Move the stdin entrypoint into main() and gate it behind require.main === module - pre-write-doc-warn.js now calls main() explicitly instead of relying on the import side effect - Add regression tests: require() attaches no stdin listeners, run()/main() stay exported, and the pre-write-doc-warn shim still warns * docs(hooks): add JSDoc for doc-file-warning main() entrypoint Satisfies the docstring-coverage pre-merge check; documents the stdin entrypoint and why it must not run on require(). * fix(windows): prefer PowerShell over bash to prevent zombie process accumulation (#2346) * fix(windows): prefer PowerShell over bash to prevent zombie process accumulation On Windows, ECC hook scripts were spawning bash.exe (MSYS2/Git Bash) on every tool use via findShellBinary(). These processes were not reaped by Windows, causing 40+ zombie bash.exe/conhost.exe processes per session with noticeable system lag. Changes to scripts/hooks/plugin-hook-bootstrap.js: - Add isPowerShellBin(bin) helper: basename-based detection so full paths like C:\Windows\...\powershell.exe are handled correctly - findShellBinary(): check BASH env var first (preserves escape hatch), then on win32 probe pwsh.exe -> powershell.exe -> bash.exe -> bash; use correct probe args per shell type; cache result in _cachedShell - findBashBinary(): separate cached bash-only finder used by spawnShell .sh fallback; skips PowerShell binaries even if BASH points to one - spawnShell(): use isPowerShellBin() to select -NoProfile -NonInteractive -File args for PowerShell; .sh scripts fall back to findBashBinary() with a skip-warning if no bash found on Windows observe-runner.js is intentionally unchanged: it always invokes observe.sh which is bash-only; routing it through PowerShell would silently break it. The observe.sh -> observe.js migration is tracked separately. Fixes #2345 * fix(windows): address CodeRabbit and Greptile review comments - Add timeout: 30000 to all spawnSync probe calls in findShellBinary and findBashBinary to prevent hangs on broken/stalled shell candidates - Add -ExecutionPolicy Bypass to PowerShell -File invocation to fix execution on machines with the default Restricted policy (Win10/11) - Add PowerShell availability skip guard to PS selection test (mirrors existing bash skip guard) - Fix no-bash test to keep PowerShell on PATH so the .sh fallback branch is actually exercised rather than hitting shell-unavailable early exit * test: add timeout to spawnSync probes in Windows test skip guards --------- Co-authored-by: Christopher J Diamond <[email protected]> * feat(session): LLM-powered session summary via claude -p (#2388) Replace mechanical text extraction in session-end.js and pre-compact.js with LLM-generated summaries using `claude -p`. Summaries now capture design decisions, resolved bugs, changed files, and carry-over context rather than just truncated user message snippets. - Add scripts/lib/llm-summary.js: generateSessionSummary, extractConversationText, getContextRemainingPct, getContextThreshold, getLLMModel - Update scripts/hooks/session-end.js: trigger LLM when context < 20% or every 50 messages (env-configurable via ECC_LLM_SUMMARY_*) - Update scripts/hooks/pre-compact.js: generate LLM summary right before compaction and write it to the active session .tmp file - Add tests/lib/llm-summary.test.js: 18 unit tests - Update tests/hooks/hooks.test.js: 3 integration tests for new behaviour Recursion guard: sets ECC_SKIP_LLM_SUMMARY=1 in subprocess env so Stop hooks fired by the claude -p subprocess do not re-enter summarisation. Requires no ANTHROPIC_API_KEY — reuses Claude Code's own authentication. Co-authored-by: Hiroshi Tanaka <[email protected]> Co-authored-by: Claude Sonnet 4.6 <[email protected]> * chore(deps): update anthropic requirement from >=0.25.0 to >=0.111.0 (#2329) Updates the requirements on [anthropic](https://github.com/anthropics/anthropic-sdk-python) to permit the latest version. - [Release notes](https://github.com/anthropics/anthropic-sdk-python/releases) - [Changelog](https://github.com/anthropics/anthropic-sdk-python/blob/main/CHANGELOG.md) - [Commits](https://github.com/anthropics/anthropic-sdk-python/compare/v0.25.0...v0.111.0) --- updated-dependencies: - dependency-name: anthropic dependency-version: 0.111.0 dependency-type: direct:production ... Signed-off-by: dependabot[bot] <[email protected]> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * chore(deps): bump actions/checkout from 6.0.3 to 7.0.0 (#2328) Bumps [actions/checkout](https://github.com/actions/checkout) from 6.0.3 to 7.0.0. - [Release notes](https://github.com/actions/checkout/releases) - [Commits](https://github.com/actions/checkout/compare/v6.0.3...v7) --- updated-dependencies: - dependency-name: actions/checkout dependency-version: 7.0.0 dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <[email protected]> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * chore(deps): bump slsa-framework/slsa-github-generator/.github/workflows/generator_generic_slsa3.yml (#2330) Bumps [slsa-framework/slsa-github-generator/.github/workflows/generator_generic_slsa3.yml](https://github.com/slsa-framework/slsa-github-generator) from 1.4.0 to 2.1.0. - [Release notes](https://github.com/slsa-framework/slsa-github-generator/releases) - [Changelog](https://github.com/slsa-framework/slsa-github-generator/blob/main/CHANGELOG.md) - [Commits](https://github.com/slsa-framework/slsa-github-generator/compare/68bad40844440577b33778c9f29077a3388838e9...f7dd8c54c2067bafc12ca7a55595d5ee9b75204a) --- updated-dependencies: - dependency-name: slsa-framework/slsa-github-generator/.github/workflows/generator_generic_slsa3.yml dependency-version: 2.1.0 dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <[email protected]> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * chore(deps): bump cron from 0.16.0 to 0.17.0 in /ecc2 (#2333) Bumps [cron](https://github.com/zslayton/cron) from 0.16.0 to 0.17.0. - [Release notes](https://github.com/zslayton/cron/releases) - [Commits](https://github.com/zslayton/cron/commits) --- updated-dependencies: - dependency-name: cron dependency-version: 0.17.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <[email protected]> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * chore(deps-dev): update pytest requirement from >=8.0 to >=9.1.1 (#2324) Updates the requirements on [pytest](https://github.com/pytest-dev/pytest) to permit the latest version. - [Release notes](https://github.com/pytest-dev/pytest/releases) - [Changelog](https://github.com/pytest-dev/pytest/blob/main/CHANGELOG.rst) - [Commits](https://github.com/pytest-dev/pytest/compare/8.0.0...9.1.1) --- updated-dependencies: - dependency-name: pytest dependency-version: 9.1.1 dependency-type: direct:development ... Signed-off-by: dependabot[bot] <[email protected]> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * chore(deps-dev): update mypy requirement from >=1.10 to >=2.1.0 (#2326) Updates the requirements on [mypy](https://github.com/python/mypy) to permit the latest version. - [Changelog](https://github.com/python/mypy/blob/master/CHANGELOG.md) - [Commits](https://github.com/python/mypy/compare/v1.10.0...v2.1.0) --- updated-dependencies: - dependency-name: mypy dependency-version: 2.1.0 dependency-type: direct:development ... Signed-off-by: dependabot[bot] <[email protected]> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * chore(deps-dev): update pytest-cov requirement from >=4.1 to >=7.1.0 (#2332) Updates the requirements on [pytest-cov](https://github.com/pytest-dev/pytest-cov) to permit the latest version. - [Changelog](https://github.com/pytest-dev/pytest-cov/blob/master/CHANGELOG.rst) - [Commits](https://github.com/pytest-dev/pytest-cov/compare/v4.1.0...v7.1.0) --- updated-dependencies: - dependency-name: pytest-cov dependency-version: 7.1.0 dependency-type: direct:development ... Signed-off-by: dependabot[bot] <[email protected]> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * chore(deps): bump the actions-minor-and-patch group across 1 directory with 3 updates (#2325) Bumps the actions-minor-and-patch group with 3 updates in the / directory: [actions/setup-node](https://github.com/actions/setup-node), [pnpm/action-setup](https://github.com/pnpm/action-setup) and [softprops/action-gh-release](https://github.com/softprops/action-gh-release). Updates `actions/setup-node` from 6.3.0 to 6.4.0 - [Release notes](https://github.com/actions/setup-node/releases) - [Commits](https://github.com/actions/setup-node/compare/v6.3.0...48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e) Updates `pnpm/action-setup` from 6.0.8 to 6.0.9 - [Release notes](https://github.com/pnpm/action-setup/releases) - [Commits](https://github.com/pnpm/action-setup/compare/0e279bb959325dab635dd2c09392533439d90093...0ebf47130e4866e96fce0953f49152a61190b271) Updates `softprops/action-gh-release` from 3.0.0 to 3.0.1 - [Release notes](https://github.com/softprops/action-gh-release/releases) - [Changelog](https://github.com/softprops/action-gh-release/blob/master/CHANGELOG.md) - [Commits](https://github.com/softprops/action-gh-release/compare/b4309332981a82ec1c5618f44dd2e27cc8bfbfda...718ea10b132b3b2eba29c1007bb80653f286566b) --- updated-dependencies: - dependency-name: actions/setup-node dependency-version: 6.4.0 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: actions-minor-and-patch - dependency-name: pnpm/action-setup dependency-version: 6.0.9 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: actions-minor-and-patch - dependency-name: softprops/action-gh-release dependency-version: 3.0.1 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: actions-minor-and-patch ... Signed-off-by: dependabot[bot] <[email protected]> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * chore(deps): bump the cargo-minor-and-patch group across 1 directory with 3 updates (#2387) Bumps the cargo-minor-and-patch group with 3 updates in the /ecc2 directory: [ratatui](https://github.com/ratatui/ratatui), [anyhow](https://github.com/dtolnay/anyhow) and [uuid](https://github.com/uuid-rs/uuid). Updates `ratatui` from 0.30.1 to 0.30.2 - [Release notes](https://github.com/ratatui/ratatui/releases) - [Changelog](https://github.com/ratatui/ratatui/blob/main/CHANGELOG.md) - [Commits](https://github.com/ratatui/ratatui/compare/ratatui-v0.30.1...ratatui-v0.30.2) Updates `anyhow` from 1.0.102 to 1.0.103 - [Release notes](https://github.com/dtolnay/anyhow/releases) - [Commits](https://github.com/dtolnay/anyhow/compare/1.0.102...1.0.103) Updates `uuid` from 1.23.3 to 1.23.4 - [Release notes](https://github.com/uuid-rs/uuid/releases) - [Commits](https://github.com/uuid-rs/uuid/compare/v1.23.3...v1.23.4) --- updated-dependencies: - dependency-name: ratatui dependency-version: 0.30.2 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: cargo-minor-and-patch - dependency-name: anyhow dependency-version: 1.0.103 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: cargo-minor-and-patch - dependency-name: uuid dependency-version: 1.23.4 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: cargo-minor-and-patch ... Signed-off-by: dependabot[bot] <[email protected]> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * chore(deps-dev): bump eslint from 9.39.2 to 10.6.0 (#2260) Bumps [eslint](https://github.com/eslint/eslint) from 9.39.2 to 10.6.0. - [Release notes](https://github.com/eslint/eslint/releases) - [Commits](https://github.com/eslint/eslint/compare/v9.39.2...v10.6.0) --- updated-dependencies: - dependency-name: eslint dependency-version: 10.5.0 dependency-type: direct:development update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <[email protected]> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * fix(ci): unbreak main after dependabot batch (checkout SHA + lint) (#2393) * fix(ci): track actions/checkout v7 SHA in supply-chain workflow test Dependabot #2328 bumped actions/checkout v6->v7, changing the pinned SHA in supply-chain-watch.yml; update the test's expected SHA to match. * Revert "feat(workflows): add orch-review native Workflow pilot (#2363)" This reverts commit 1031d312ccd16925c903174293b4cec4ecba1001. * feat: add ecc-recipes skill (#2319) * feat: add ecc-recipes skill Maps a described workflow to the right ECC command-group with run-order and stop condition, and browses command-group recipe families. Fills the gap between ecc-guide (flat catalog) and prompt-optimizer (single-prompt match) by adding family grouping, run-order, and stop conditions. Advisory only; reads commands/ live. * fix(ecc-recipes): address review - flatten frontmatter origin/author/version to top-level (repo convention) - guard unset CMD_DIR before globbing; use find instead of ls - show burn-warning explicitly in output template * feat(ecc-recipes): add argument-hint for slash UI * feat(skills): add mailtrap-email-integration skill (#2288) Adds a new Tool Integration skill (mailtrap-email-integration) covering transactional email sending patterns: sandbox vs. production separation, API authentication, and domain verification. Focused on patterns that generalize beyond one vendor, per the repo's Skill Adaptation Policy. * docs(code-tour): document ref-field semantics to prevent PR-tour file-not-found (#2273) The code-tour skill mentioned the CodeTour 'ref' field only in an example, with no explanation of its behavior. CodeTour resolves each step's file content from the git revision named by 'ref' (not the working tree) whenever ref differs from HEAD, so any file that does not exist at that revision fails to open with 'The editor could not be opened because the file was not found' - even though the file is present on disk. This bit a generated PR tour where ref was set to the base branch (develop): every file ADDED by the PR is absent on the base, so all new-file steps 404'd while the tour tree and comments still rendered, making the cause non-obvious. Adds a 'The ref Field' section explaining the resolution behavior and the rule that PR tours must pin ref to the branch head (never the base), plus a validation step to confirm every referenced file exists at the chosen ref. * fix(gateguard): finish tool-agnostic checklist across edit gate and SKILL.md copies (#2274) b3268fef (#2272) made the write-gate "confirm no existing file" item tool-agnostic in the JS hook, but the rest of the checklist surface still names Glob/Grep. On hosts without those tools the agent still hits a dead tool call on: - the edit-gate "list importers" item in the hook (scripts/hooks/gateguard-fact-force.js) - both checklist items in all three SKILL.md copies (en, ja-JP, zh-CN) Apply the same wording b3268fef introduced — "(search the tree — Glob/Grep, or find/grep via Bash)" — to those five remaining spots so the whole gate is consistent. Prose-only; no logic change. Follow-up to #2272 / b3268fef. * feat(skills): harden the file upload validation section in django-security (#2338) * feat(skills): harden the file upload validation section in django-security * Update skills/django-security/SKILL.md Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> * add missing stuff to second code block * add import to the top of the code block --------- Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> * docs(skills): update Prisma and Zod API patterns for cross-version compatibility (#2336) * docs(skills): update Prisma and Zod API patterns for cross-version compatibility - skills/prisma-patterns: show both adapter-based and direct PrismaClient initialization side-by-side; update import paths with conditional notes; rewrite version header to be release-agnostic - skills/backend-patterns: fix ZodError.errors -> ZodError.issues - skills/coding-standards: fix ZodError.errors -> ZodError.issues - skills/security-review: fix ZodError.errors -> ZodError.issues These API differences were discovered during implementation of a full-stack health assessment project. The updated code samples show both the new and old API forms so the skill remains useful regardless of which Prisma or Zod version is installed. Closes #2335 * fix(skills): revert Prisma client imports to '@prisma/client' The 'prisma' npm package is the CLI tool, not the runtime client. Using it as an import source would cause compile-time failures on all versions. '@prisma/client' remains the correct import source for the generated PrismaClient and Prisma namespace types. Found by Greptile during PR review. * feat(skills): make tdd-workflow test-runner aware (npm/pnpm/yarn/bun) (#2347) * feat(skills): make tdd-workflow test-runner aware (npm/pnpm/yarn/bun) Add "Step 0: Detect the Test Runner" so the RED/GREEN cycle no longer hardcodes `npm test`. Distinguishes the package manager from the test runner (a project can install with Bun yet run Jest/Vitest), adds a runner command matrix, and warns about `bun test` (native bun:test runner) vs `bun run test` (runs the package.json script) — a common ESM failure mode. Adds a Bun native test pattern section and links the bun-runtime skill. Applied to both the canonical skills/ copy and the .agents/skills/ Codex subset (manual sync per CONTRIBUTING). * docs(skills): apply <test>/<coverage> placeholders in tdd-workflow steps Address review feedback on PR #2347: Step 0 instructs the agent to substitute the detected runner command, but Steps 3/5/7, Run Coverage Report, Watch Mode, Pre-Commit, and CI/CD still showed literal `npm test` / `npm run test:coverage` — so an agent reaching those blocks could run npm test on a pnpm/bun project. Replace them with the <test> / <test-watch> / <coverage> placeholders from Step 0. Left untouched: the plan-handoff allowlist example and the Step 8 evidence-table samples (illustrative, not run-this instructions). Applied to both the canonical and Codex-subset copies. * docs(skills): make pre-commit lint runner-agnostic via <lint> placeholder Follow-up to PR #2347 review (CodeRabbit): the pre-commit example still used `npm run lint`, coupling it to npm after test/coverage were made runner-aware. Add a `<lint>` column to the Step 0 runner matrix (npm run lint / pnpm lint / yarn lint / bun run lint) and change the Pre-Commit Hook example to `<test> && <lint>`. Applied to both the canonical and Codex-subset copies. * chore: re-trigger CI (flaky windows/node20 npm cell) * refactor(commands): remove duplicated content in skill-create and learn-eval (#2348) skill-create: drop the "Example Output" section (53 lines) — it re-rendered the same skeleton already defined by the Step 3 output template, just with filled-in `my-app` values. learn-eval: drop the "Next Action" column from the 5b verdict table — it duplicated Step 6's "Verdict-specific confirmation flow". The table now carries Verdict + Meaning, and a pointer to Step 6 as the single source for each verdict's action. No behavior, frontmatter, or design-rationale changes. * chore(catalog): sync manifests after skill batch (#2319 #2288 #2273 #2274 #2338 #2336 #2347 #2348) (#2394) Regenerate catalog doc counts + command registry after merging the verified skill/agent batch. Local full suite was green (2924/2924) with these applied. * fix(clv2): align Python _update_registry schema with shell counterpart (#2369) * fix(clv2): align Python _update_registry schema with shell counterpart The Python `_update_registry` in instinct-cli.py wrote registry entries without the `id` and `created_at` fields, while the shell counterpart in detect-project.sh writes both. A projects.json entry could therefore have a different shape depending on which path (Python CLI or shell hook) last touched it. Emit the same field set and order as the shell version: id, name, root, remote, created_at (preserved from any existing entry), last_seen. Add regression tests asserting field parity and created_at preservation. Fixes #2299 * fix(clv2): guard _update_registry against a non-dict registry entry A malformed projects.json (a non-dict value for the current project id, e.g. null) would make existing.get("created_at", ...) raise and crash the update, losing the old code's ability to self-heal a corrupt per-entry value. Normalize existing to {} when it is not a dict so the entry is healed by the rewrite. Add a regression test for the malformed-entry path. * test(clv2): assert the first-write created_at == last_seen contract The new _update_registry tests only checked both timestamps were truthy. On the initial write both derive from the same `now`, so created_at must equal last_seen; assert that explicitly so a later refactor that breaks the contract is caught. Split the compound assertions into single-expression checks. * fix(clv2): heal a non-dict top-level registry in _update_registry A projects.json that is valid JSON but not a mapping (e.g. `[]` or a string) previously crashed _update_registry on registry.get(), before the per-entry guard could run, so the corrupt file could not be healed. Guard the top-level shape right after the load and fall back to {} so the rewrite repairs the file — matching the per-entry healing already in place. Resolves the remaining CodeRabbit finding on #2299. Co-Authored-By: Claude Opus 4.8 <[email protected]> --------- Co-authored-by: Claude Opus 4.8 <[email protected]> * fix(clv2): serialize observer signal-counter to stop dropped increments (#2372) observe.sh bumps the SIGUSR1 throttle counter in ${PROJECT_DIR}/.observer-signal-counter with an unlocked read-modify-write. The hook runs on every tool call, so concurrent invocations read the same value, both increment, and lose a write, signaling the observer at unpredictable intervals and defeating the #521 throttle. Serialize the read-modify-write under a lock, and only ever bump the counter while that lock is held: - Prefer flock with a bounded -w wait (the OS auto-releases it when the fd closes or the process dies, so there is no stale lock and no lost increment); on a timeout the tick is skipped rather than bumped unlocked. - Fall back to an atomic mkdir lock on platforms without flock, with a bounded spin. An EXIT trap cleans up on normal completion; INT/TERM traps release the lock and exit, so a signal cannot drop the lock and then continue the read-modify-write without ownership. If the lock cannot be acquired in the budget the tick is skipped rather than raced. No hand-rolled PID stale-reclaim (which is racy and can delete a live re-acquirer's lock). - Guard the counter read against a corrupt (non-integer) file that would abort the hook under set -e. Add tests/hooks/observe-signal-counter-race.test.js: 20 concurrent observe.sh invocations must not lose increments (exact under flock; at most one dropped on the best-effort mkdir fallback), the runner rejects on any hook execution failure or hang, plus content guards for the lock and the corrupt-counter handling. Fixes #2296 * fix(clv2): surface SIGALRM timeout drops in observe.sh (#2373) * fix(clv2): surface SIGALRM timeout drops in observe.sh The inline-Python observation writers in observe.sh arm a signal.SIGALRM alarm (8s) so they self-terminate before the async hook's 10s timeout can orphan them (#2278). The handler _ecc_bail called sys.exit(0) with no logging, so when the alarm fired the in-flight observation was silently dropped: nothing was logged, no partial write occurred, and the shell saw a clean exit. There was no way to detect or count how many observations were being lost. Add a single stderr visibility line to both _ecc_bail handlers (the parse-error fallback path and the main observation-writing path) before sys.exit(0), using the repo's "[observe]" log prefix. Exit code stays 0: in a Claude Code hook a non-zero exit signals a block, so changing it would turn an internal timeout into a user-facing tool block. The warning goes to stderr (not stdout) because both blocks redirect stdout into the observations file. Add tests/hooks/observe-signal-timeout.test.js: a static regression guard that every _ecc_bail handler logs to stderr before exiting and keeps exit 0, plus a behavioral check that runs the real handler text extracted from observe.sh and confirms a fired alarm exits 0 and emits the [observe] warning on stderr only. Fixes #2300 * test(clv2): exercise both _ecc_bail handlers end-to-end The behavioral SIGALRM-fire test ran only handlers[0] (the parse-error fallback path); the main observation-write path (handlers[1]) was covered only by the static regex guard. The write path is the higher-value one to verify end-to-end since it carries valid, parseable data that would succeed given more time, so a silent drop there is the worst case. Loop the behavioral check over every extracted handler so a regression that silenced the second handler's stderr write is caught at runtime, not just by the static guard. * test(clv2): select timeout handlers by marker, not array index The behavioral check looped over all extracted _ecc_bail handlers by index. If an unrelated _ecc_bail were ever added to observe.sh, the loop would either test the wrong block or be diluted. Filter the handlers to those carrying the "[observe] SIGALRM timeout" marker so the live SIGALRM check stays pinned to the two #2300 timeout handlers regardless of array order or future additions. * test(clv2): fail fast when python is missing in SIGALRM check The behavioral test returned early when no python interpreter was found, which the test harness records as a PASS — so the SIGALRM contract could go entirely unverified yet still look green. Throw instead, matching the existing insaits-security-monitor convention of failing when a required Python runtime is absent, and drop the in-test console.log. * test(clv2): add coverage for instinct-cli prune, projects ops, promote dry-run, normalize-url (#2374) * test(clv2): cover instinct-cli prune, projects ops, promote dry-run, normalize-url Add pytest coverage for previously-untested functions in skills/continuous-learning-v2/scripts/instinct-cli.py: - _normalize_remote_url: scp/https/file forms, credential + .git stripping, network lowercasing, case-preserving local paths, idempotence - _promote_specific dry-run: returns 0 and writes no global file - projects delete/gc/merge: invalid-id, not-found, dry-run, and force paths over registry + storage, asserting destructive ops are gated - cmd_prune: dry-run keeps files; non-dry-run deletes only expired; quiet Test-only change; no production code modified. Fixes #2302 * test(clv2): assert dry-run storage no-op and quiet-mode stderr silence Address CodeRabbit review on #2374: - projects gc/merge dry-run tests now also assert on-disk storage is untouched (empty1 project dir survives; nothing copied into dest personal), closing the gap where a storage-mutating dry-run regression would still pass. - cmd_prune quiet test now asserts stderr is empty too, not just stdout. * test(clv2): cover merge missing-destination and prune empty-pending branches * fix(clv2): archive observations only after successful analysis in observer-loop (#2386) analyze_observations moved observations.jsonl into observations.archive/ unconditionally, even when the Claude analysis failed (timeout, non-zero exit, rate limit). Because the analyzer only reads the live file, a failed batch was archived and never re-analyzed, silently dropping the instincts it would have produced. Return early on a non-zero analysis exit so the archive mv runs only on success, retaining observations for the next cycle to retry. Resolve the script's own directory from ${BASH_SOURCE[0]} (SCRIPT_DIR) so sibling scripts (session-guardian.sh) and relative helpers resolve correctly under both execution and sourcing, and add a source-guard so observer-loop.sh can be sourced without starting the loop. Add a regression test covering both the failure (retain) and success (archive) paths. Fixes #2370 * feat(continuous-learning-v2): make observer model configurable via ECC_OBSERVER_MODEL (#2390) * feat(continuous-learning-v2): make observer model configurable via ECC_OBSERVER_MODEL The observer hardcoded `--model haiku`. Parameterize as "${ECC_OBSERVER_MODEL:-haiku}": the haiku default is preserved (no behavior change for existing users), but users can opt into a stronger model — e.g. `ECC_OBSERVER_MODEL=opus` — for higher-quality instinct extraction. Useful on subscription plans where model cost isn't the limiting factor. * fix(continuous-learning-v2): address review — update wiring test + docs - Update source-inspection test to assert the ${ECC_OBSERVER_MODEL:-haiku} defaulting behavior (was matching the literal `claude --model haiku`, which this PR changed). All 31 tests pass. - Add guidance to raise ECC_OBSERVER_TIMEOUT_SECONDS for slower models (e.g. opus) so the 120s watchdog doesn't kill analysis mid-run. - Fix now-stale 'Haiku session' comment -> 'observer session' (model is configurable). * feat(rules,skills): add React Native / Expo rules pack and react-native-patterns skill (#2275) * feat(rules,skills): add React Native / Expo rules pack and react-native-patterns skill * fix(rules,skills): address review feedback — safeParse nav example, drop deprecated sentry-expo, memoize list renderItem, clarify New Architecture SDK support * fix(rules,skills): drop deprecated Flipper, surface permission-denied state in location hook * Add growth-log skill: methodology for effective learning capture (#2377) * Add growth-log skill: methodology for writing effective, transferable growth log entries * Add metadata.origin: ECC frontmatter per repo convention (Greptile feedback) * Re-sign: apply GPG-verified commit to growth-log branch (rebase artifact, content unchanged) * docs(growth-log): v1.1.0 — remove personal library structure, generic storage, delivery-gate optional companion * Stop hook: verify thinking quality at session end — task completeness, assumptions, stale logs, disk space (delivery-gate) (#2378) * Restore delivery-gate: Stop hook with learning capture enforcement (auto-closed by fork sync, now on clean branch) * Fix bot findings: log level→INFO (DISK_REMIND dead code), count_edits full transcript (not truncated), memory-dir-absent warning (not silent pass), SKILL.md description accuracy * Fix CodeRabbit feedback: treat missing memory-dir as all-stale on complex tasks (fail-close instead of fail-open) * Trigger bot re-review (no logic changes) * Fix: handle both stdin formats — raw transcript AND JSON with transcript_path (Greptile feedback) * Add debug log for memory-dir lookup path * Fix path encoding: replace colon with dash (not strip), matching Claude Code actual encoding on Windows * Fix SKILL.md: update How It Works for JSON+transcript_path, add English translation to CLAUDE.md block (Greptile feedback) * Fix: memory-dir absent → warn but don't block (prevents deadlock for new users per Greptile feedback) * fix: restore daltino-approved voice (thinking quality/收尾铁律) with technical patches Reverts 'session hygiene' rebranding. Preserves original approved framing while keeping technical improvements: - JSON transcript_path parsing documentation - filesystem mtime staleness check - 'skip tests for now' rationalization pattern - disk critically low explicit block condition * fix: remove stdout JSON echo — Stop hooks write feedback to stderr, not stdout Previously sys.stdout.write(raw) echoed the raw hook JSON payload to stdout, which Claude Code displays as the hook's response message. When the hook blocked (exit 2), Claude saw {"transcript_path":"...","session_id":"..."} instead of the actual blocking reason from stderr. This made the gate functionally silent from Claude's perspective — it could not guide Claude to the corrective action (update growth-log / free disk). Fix per Greptile feedback: stop echo, let stderr messages reach Claude. * fix: remove duplicate disk-critical log line * docs(delivery-gate): v1.1.0 — accurate scope (deterministic checks, not reasoning), warning vs block table, CI/CD analogy, limitations section, self-audit pairing * fix(delivery-gate): expand rationalization regex coverage (R3/R4) — match "we can fix" and "integration tests" variants * chore: bump version to 1.1.1 to re-trigger CI checks * feat: add loop-design-check skill (design + review goal-oriented agent loops) (#2381) * chore(catalog): sync manifests + fix skill emoji (wave 2) (#2395) * chore(catalog): sync manifests after skill batch (#2275 #2377 #2378 #2381) Update skill counts (273 -> 277) across catalog docs after the verified skill batch. * fix(skills): replace emoji with ASCII in growth-log + loop-design-check check-unicode-safety (pre-push gate) bans emoji in SKILL.md; the merged #2377 and #2381 slipped through run-all.js. Swap U+274C/U+2705 for 'Avoid:'/'Bad:'/'Good:'. * fix(plan-orchestrate): detect ecc@ecc marketplace + emit ecc: agent prefix (#2316) (#2409) * fix(plan-orchestrate): detect ecc@ecc marketplace + emit ecc: agent prefix (#2316) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(ci): resync lockfiles with package.json (eslint 10) + migrate yarn.lock to Yarn 4 format package.json requires eslint@^10.6.0 but the committed locks pinned 9.39.2, so npm ci aborted and Yarn 4 hardened mode rejected the stale v1-classic yarn.lock (YN0028). Regenerate package-lock.json and rewrite yarn.lock in Yarn 4 (berry) format so npm ci and immutable yarn installs both pass. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(ci): require clean probe exit for Windows shell/bash detection; add pyyaml dev dep Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: affaan <[email protected]> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor: consolidate duplicated hook-root resolver into shared resolveEccRoot() (#2368) (#2410) * fix(ci): resync lockfiles with package.json (eslint 10) + migrate yarn.lock to Yarn 4 format package.json requires eslint@^10.6.0 but the committed locks pinned 9.39.2, so npm ci aborted and Yarn 4 hardened mode rejected the stale v1-classic yarn.lock (YN0028). Regenerate package-lock.json and rewrite yarn.lock in Yarn 4 (berry) format so npm ci and immutable yarn installs both pass. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(ci): require clean probe exit for Windows shell/bash detection; add pyyaml dev dep Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor: consolidate duplicated hook-root resolver into shared resolveEccRoot() (#2368) The inline node -e resolver blob was duplicated ~60x across hooks.json, command docs, and translations. Each copy inlined the full ~700-char plugin-root search using a spread over nested array literals (p.join(d,'plugins',...s) over [['ecc'],...]), which breaks Windows hook execution due to shell quoting (#2368). Collapse every copy to a 250-char locator that loads the committed resolve-ecc-root module and delegates to resolveEccRoot() — no spread, no nested array literals, no escaped double quotes. The real search logic now lives in one tested module. Also route session-start-bootstrap.js through resolveEccRoot() instead of its own duplicated reimplementation, and fix the auto-update.md 'marketplace' (singular) typo along the way. Guard tests updated: discovery behavior is asserted against resolveEccRoot(); the inline is asserted to delegate and to contain no Windows-fragile constructs. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(resolve-ecc-root): restore full env-unset discovery in inline resolver Address Greptile review on #2410: when CLAUDE_PLUGIN_ROOT is unset the delegating inline could only load the resolver module from ~/.claude, returning ~/.claude without ever reaching the plugin/cache search. Restore the old inline's discovery breadth (exact plugin roots + versioned cache) Windows-safely (no spread, nested arrays, or escaped quotes), then delegate the authoritative decision to resolveEccRoot(). Add regression tests for plugin-subdir and versioned-cache bootstrap with env unset. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: affaan <[email protected]> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix: docs/COMMAND-REGISTRY.json check fails on fresh Windows clone (missing .gitattributes) (#2437) * fix: add .gitattributes to force LF line endings for text files npm run command-registry:check (part of npm test) fails on a fresh clone on Windows with the common core.autocrlf=true setting: git checks out docs/COMMAND-REGISTRY.json with CRLF, but generate-command-registry.js always writes LF, so the strict string comparison in checkRegistry() never matches. Forcing LF via .gitattributes makes checkouts consistent across platforms regardless of a contributor's local autocrlf setting. * fix: normalize CRLF line endings to LF per .gitattributes pyproject.toml, src/llm/__init__.py, src/llm/prompt/builder.py, src/llm/providers/claude.py, and tests/test_builder.py had CRLF line endings committed to the repo, inconsistent with the rest of the codebase. Renormalized via 'git add --renormalize .' now that .gitattributes enforces eol=lf. --------- Co-authored-by: Affaan Mustafa <[email protected]> * feat(workflows): re-land orch-review workflow + add /orch-review command (#2400) * feat(workflows): re-land orch-review workflow + add /orch-review command Re-lands #2363 (reverted by #2393 to unbreak main's lint) and fixes the root cause so it stays green: - Restore workflows/orch-review.workflow.js + workflows/README.md. - eslint.config.js: ignore 'workflows/**/*.workflow.*' and '.claude/workflows/**' per the maintainer's note in #2393. Workflow DSL scripts use both top-level export (ESM) and top-level return (the runtime wraps them in an async fn), which no single eslint sourceType can parse — they must be excluded, not lint-fixed. 'npx eslint .' is green with this ignore. - Add commands/orch-review.md (the /orch-review surface) + regenerate docs/COMMAND-REGISTRY.json. Supersedes #2397 (command-only), which referenced the reverted workflow. * fix(workflows): address orch-review bot review findings - Verifier uncertainty no longer demotes blockers (Greptile P1 + CodeRabbit): isReal=false only refutes when confidence >= 0.8; low-confidence 'false' is treated as uncertain and kept blocking (fail closed). - Treat the diff (and finding text) as untrusted input in both review and verify prompts; ignore embedded directives (prompt-injection hardening). - Validate changedFiles entries are strings, not just that it is an array. - Enforce proof for HIGH/CRITICAL in FINDINGS_SCHEMA, not only in the prompt. - Remove in-place mutation in dimension build + dedup merge (immutable). - /orch-review: extract & validate a numeric PR id before shelling out to gh. - Docs: complete the stats example, soften wording, refresh follow-up list. * style(workflows): apply formatter to orch-review assembly * fix(plan-orchestrate): detect ecc@ecc marketplace + emit ecc: agent prefix (#2316) (#2409) * fix(plan-orchestrate): detect ecc@ecc marketplace + emit ecc: agent prefix (#2316) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(ci): resync lockfiles with package.json (eslint 10) + migrate yarn.lock to Yarn 4 format package.json requires eslint@^10.6.0 but the committed locks pinned 9.39.2, so npm ci aborted and Yarn 4 hardened mode rejected the stale v1-classic yarn.lock (YN0028). Regenerate package-lock.json and rewrite yarn.lock in Yarn 4 (berry) format so npm ci and immutable yarn installs both pass. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(ci): require clean probe exit for Windows shell/bash detection; add pyyaml dev dep Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: affaan <[email protected]> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor: consolidate duplicated hook-root resolver into shared resolveEccRoot() (#2368) (#2410) * fix(ci): resync lockfiles with package.json (eslint 10) + migrate yarn.lock to Yarn 4 format package.json requires eslint@^10.6.0 but the committed locks pinned 9.39.2, so npm ci aborted and Yarn 4 hardened mode rejected the stale v1-classic yarn.lock (YN0028). Regenerate package-lock.json and rewrite yarn.lock in Yarn 4 (berry) format so npm ci and immutable yarn installs both pass. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(ci): require clean probe exit for Windows shell/bash detection; add pyyaml dev dep Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor: consolidate duplicated hook-root resolver into shared resolveEccRoot() (#2368) The inline node -e resolver blob was duplicated ~60x across hooks.json, command docs, and translations. Each copy inlined the full ~700-char plugin-root search using a spread over nested array literals (p.join(d,'plugins',...s) over [['ecc'],...]), which breaks Windows hook execution due to shell quoting (#2368). Collapse every copy to a 250-char locator that loads the committed resolve-ecc-root module and delegates to resolveEccRoot() — no spread, no nested array literals, no escaped double quotes. The real search logic now lives in one tested module. Also route session-start-bootstrap.js through resolveEccRoot() instead of its own duplicated reimplementation, and fix the auto-update.md 'marketplace' (singular) typo along the way. Guard tests updated: discovery behavior is asserted against resolveEccRoot(); the inline is asserted to delegate and to contain no Windows-fragile constructs. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(resolve-ecc-root): restore full env-unset discovery in inline resolver Address Greptile review on #2410: when CLAUDE_PLUGIN_ROOT is unset the delegating inline could only load the resolver module from ~/.claude, returning ~/.claude without ever reaching the plugin/cache search. Restore the old inline's discovery breadth (exact plugin roots + versioned cache) Windows-safely (no spread, nested arrays, or escaped quotes), then delegate the authoritative decision to resolveEccRoot(). Add regression tests for plugin-subdir and versioned-cache bootstrap with env unset. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: affaan <[email protected]> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix: docs/COMMAND-REGISTRY.json check fails on fresh Windows clone (missing .gitattributes) (#2437) * fix: add .gitattributes to force LF line endings for text files npm run command-registry:check (part of npm test) fails on a fresh clone on Windows with the common core.autocrlf=true setting: git checks out docs/COMMAND-REGISTRY.json with CRLF, but generate-command-registry.js always writes LF, so the strict string comparison in checkRegistry() never matches. Forcing LF via .gitattributes makes checkouts consistent across platforms regardless of a contributor's local autocrlf setting. * fix: normalize CRLF line endings to LF per .gitattributes pyproject.toml, src/llm/__init__.py, src/llm/prompt/builder.py, src/llm/providers/claude.py, and tests/test_builder.py had CRLF line endings committed to the repo, inconsistent with the rest of the codebase. Renormalized via 'git add --renormalize .' now that .gitattributes enforces eol=lf. --------- Co-authored-by: Affaan Mustafa <[email protected]> * chore(catalog): sync command counts (92->93) + register orch-review in agent.yaml surface --------- Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: affaan <[email protected]> Co-authored-by: Boube <[email protected]> Co-authored-by: Affaan Mustafa <[email protected]> * docs: replace personal absolute paths with repo-relative agentshield/ paths --------- Signed-off-by: dependabot[bot] <[email protected]> Co-authored-by: Gaurav Dubey <[email protected]> Co-authored-by: JongHyeok Park <[email protected]> Co-authored-by: Daniel Nguyen <[email protected]> Co-authored-by: Gabriel Pitrella <[email protected]> Co-authored-by: Angad Singh Thind <[email protected]> Co-authored-by: angadsingh7666 <[email protected]> Co-authored-by: quadcent <[email protected]> Co-authored-by: Vladimir Đuranović <[email protected]> Co-authored-by: Cursor <[email protected]> Co-authored-by: Yang Cheng <[email protected]> Co-authored-by: jan <[email protected]> Co-authored-by: Tahiti18 <[email protected]> Co-authored-by: Claude Sonnet 4.6 <[email protected]> Co-authored-by: phobicdotno <[email protected]> Co-authored-by: Karstein Phobic Nyvold Kvistad <[email protected]> Co-authored-by: SSH._.WORLD <[email protected]> Co-authored-by: ChrisD <[email protected]> Co-authored-by: Christopher J Diamond <[email protected]> Co-authored-by: Hiroshi Tanaka <[email protected]> Co-authored-by: Hiroshi Tanaka <[email protected]> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: KyawZinLatt <[email protected]> Co-authored-by: Awa Dieudonne <[email protected]> Co-authored-by: Carlos Carvallo <[email protected]> Co-authored-by: Jun <[email protected]> Co-authored-by: jvirgovic <[email protected]> Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> Co-authored-by: weizhiyuan <[email protected]> Co-authored-by: jack-finance-able <[email protected]> Co-authored-by: Yeris Rifan <[email protected]> Co-authored-by: YuhaoLin2005 <[email protected]> Co-authored-by: Seekers2001 <[email protected]> Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: Boube <[email protected]>
What this is
A Stop hook that checks session hygiene before Claude can finish — not reasoning quality (that's self-audit's job), but machine-verifiable facts: file timestamps, disk usage, and regex patterns on transcript text.
Three checks, all configurable:
If a complex task (3+ file edits) completed with ALL learning libraries stale → blocks stop (exit 2). Disk critical → blocks regardless. Otherwise passes silently.
~200 lines of Python. Stdlib only. Deterministic — no AI inference, just file timestamps + regex + disk usage.
Relationship to existing ECC skills
verification-loopgateguardself-auditdelivery-gatedelivery-gate is a mechanical gate (deterministic checks on machine-verifiable facts). self-audit is a reasoning gate (agent evaluates output quality). Together they form defense in depth — the same pattern as CI pipeline gates.
Files
skills/delivery-gate/SKILL.mdskills/delivery-gate/hooks/quality-gate.pyRecent update (v1.1.0)
After studying multi-agent pipeline quality gate architectures, refined the SKILL.md documentation to accurately describe what delivery-gate actually checks (deterministic: file timestamps + regex + disk usage) and clearly separate its role from reasoning gates like self-audit. Added warning-vs-block behavior table, limitations section, and CI/CD analogy. No code changes — same hook, clearer framing.
Note
Replaces PR #2365 (auto-closed during fork sync). daltino's approval from #2365 applies to the substance — same contribution, clean feature branch.