Skip to content

[Feature][learning]: Autonomous skill creation, self-improvement, and pipeline tool calling #3683

Description

@5queezer

Summary

Add three capabilities inspired by the Hermes Agent learning loop: autonomous skill document creation from successful agent turns, self-improvement of existing skill docs during use, and a meta-tool (execute_pipeline) that collapses multi-step tool chains into a single inference call — all within ZeroClaw's single-binary, trait-driven architecture.

Problem statement

ZeroClaw currently has no mechanism to learn from successful task executions. Every session starts from the same baseline — skills must be manually authored and dropped into the workspace. Two concrete friction points:

  1. When the ZeroClawAgent solves a novel multi-step problem (e.g. a complex file + web + summarise chain), that workflow is lost. The next session repeats the same tool-call overhead.
  2. Multi-step pipelines require the LLM to reason through each tool call independently, producing multiple inference round-trips for what is semantically one task.

Hermes Agent (NousResearch) addresses this with a closed learning loop: it writes skill documents automatically, improves them during use, and collapses pipelines into single calls. ZeroClaw should match this without sacrificing its Rust single-binary guarantee or its security model.

Proposed solution

Feature A — Autonomous skill creation

After a qualifying agent turn the ZeroClawAgent silently writes a new skill
document to ~/.zeroclaw/workspace/skills/{slug}.md. From the operator's
perspective nothing extra is required — the file appears automatically and is
immediately available for future turns via the existing skill loader.

The file format matches the existing skill standard (TOML front-matter +
markdown body) so it is compatible with zeroclaw skills list, external skill
registries, and any tooling that already reads the workspace skills directory.

Config surface (additive keys in config.toml):

[skills.auto_create]
enabled = true        # default: true
max_skills = 500      # evict LRU when exceeded
similarity_threshold = 0.85  # skip if existing skill is this similar

Feature B — Skill self-improvement

When the ZeroClawAgent uses an existing skill and the task succeeds, it may
silently patch that skill's content in place. From the operator's perspective
the file is updated transparently; zeroclaw skills show {slug} reflects the
latest version and the front-matter audit trail shows what changed and when.

New CLI commands:

zeroclaw skills list              # slug · title · updated_at
zeroclaw skills show {slug}       # full file to stdout

Config surface:

[skills.auto_improve]
enabled = true          # default: true
cooldown_secs = 3600    # minimum interval between improvements per skill

Feature C — Pipeline tool calling

A new built-in tool execute_pipeline accepts a JSON payload and returns a
single aggregated result. The agent can invoke it like any other tool; no new
CLI commands or config keys are required for basic use.

Interface:

{
  "steps": [
    { "tool": "web_search", "args": { "query": "zeroclaw changelog" } },
    { "tool": "summarise",  "args": { "text": "{{step[0].result}}" } }
  ],
  "parallel": false
}

{{step[N].result}} interpolates the Nth step's output into a subsequent
step's args. Substitution is single-pass; the result is treated as an opaque
string.

The permitted tool set is declared in config so operators control the blast
radius:

[pipeline]
enabled = true
max_steps = 20
allowed_tools = ["web_search", "summarise", "read_file", "write_file"]

Non-goals / out of scope

  • No Python or Node.js runtime — single Rust binary guarantee is preserved
  • No network-based skill registry (this is local-only; ZeroMarket integration is a separate concern)
  • No changes to existing open-skills or manual skill loading behaviour
  • No modification to src/config/schema.rs public contract beyond additive keys
  • No cross-subsystem coupling: skills/ does not import channels/ internals

Alternatives considered

  1. Pre-bundled static skill document ([Feature]: Add zeroclaw skill #1948)
    Issue [Feature]: Add zeroclaw skill #1948 added a hand-authored zeroclaw.md skill so the agent has self-knowledge out of the box. That addresses discoverability but is a one-time, manually maintained file — it does not create or improve skills from experience. This proposal is the dynamic complement: skills are generated and refined at runtime, not shipped in the binary.

  2. WASM skill engine / ZeroMarket install ([Feature]: WASM skill engine — install and run tools from ZeroMarket registry #1787)
    Issue [Feature]: WASM skill engine — install and run tools from ZeroMarket registry #1787 (open) proposes fetching pre-compiled WASM tools from a registry. That is an install-time distribution mechanism — operators choose skills explicitly. The learning loop proposed here is orthogonal: it produces skills as a by-product of normal agent use, with no registry dependency. The two features can coexist; WASM-packaged skills could eventually be published from auto-generated ones.

  3. Allowing script files in skills via config ([Feature Request]: Allow script files in skills via configuration option #1889, closed)
    Issue [Feature Request]: Allow script files in skills via configuration option #1889 requested a config override for the src/skills/audit.rs script-blocking policy. That work hardened the security boundary this proposal relies on. The slug sanitisation and path assertion requirements in Features A and B are additive to the existing audit layer — not a replacement for it. No changes to src/skills/audit.rs are needed or proposed here.

  4. External self-improving-agent skill (observed in [Bug]: Zeroclaw unable to file write or edit files when prompted in Discord #2782)
    Issue [Bug]: Zeroclaw unable to file write or edit files when prompted in Discord #2782 (bug) shows a user who manually installed an external self-improving-agent skill from a third-party source. The skill was blocked by ZeroClaw's security policy. This confirms that demand exists, that the current skill security policy correctly blocks untrusted scripts, and that a native, policy-compliant implementation inside ZeroClaw is the right path rather than relying on external skills that bypass the audit layer.

Acceptance criteria

Feature A — Autonomous skill creation

  • After a qualifying agent turn (≥3 successful tool calls, non-duplicate, repeatable), a skill file is written to ~/.zeroclaw/workspace/skills/{slug}.md
  • Skill file contains valid TOML front-matter with title, tags, and created_at
  • Slug is derived from LLM output and passes ^[a-z0-9_-]{1,64}$ validation before any filesystem operation
  • Resolved skill path is asserted to be inside skills_dir; any path outside aborts with an error log
  • Skill is indexed in the FTS5 table immediately after write
  • Near-duplicate turns (cosine similarity > 0.85 to existing skill) do not produce a new file
  • When max_skills is reached, the least-recently-used skill is evicted and the eviction is logged at WARN
  • Feature is fully disabled at compile time with --no-default-features (no runtime impact)
  • All tests in the Tests required section for Feature A pass under cargo test
  • fuzz_target!(skill_create_slug_input) compiles and runs under cargo fuzz

Feature B — Skill self-improvement

  • After the agent uses an existing skill and the task completes, should_improve_skill is called
  • Improved content is written atomically: temp file → validate → rename; original is never overwritten if validation fails
  • Validation rejects empty content, non-UTF-8 bytes, and malformed TOML front-matter
  • Every successful write appends updated_at and improvement_reason to front-matter; prior audit entries are preserved
  • A skill cannot be improved more than once within the configured cooldown window (default 1 h); second call within window returns Ok(None) silently
  • zeroclaw skills list prints slug, title, and updated_at for all indexed skills
  • zeroclaw skills show {slug} prints the full skill file to stdout
  • All tests in the Tests required section for Feature B pass under cargo test

Feature C — Pipeline tool calling

  • execute_pipeline is registered as a tool and callable from the agent loop
  • Steps execute sequentially by default; parallel: true runs independent steps concurrently
  • {{step[N].result}} interpolation is single-pass; substituted values containing {{ are stripped before use
  • Any step referencing a tool not on the configured allowlist is rejected with PipelineError::UnknownTool before execution starts
  • Pipelines exceeding MAX_STEPS (default 20) are rejected with PipelineError::TooManySteps before execution starts
  • Network targets in tool args are validated against the existing ZeroClaw allowlist/blocklist; loopback, link-local, and RFC-1918 ranges return an error
  • Feature is fully disabled at compile time with --no-default-features (no runtime impact)
  • All tests in the Tests required section for Feature C pass under cargo test
  • fuzz_target!(pipeline_template_interpolate) compiles and runs under cargo fuzz

General

  • cargo build --release --locked succeeds with and without --no-default-features
  • cargo clippy -- -D warnings passes on all new modules
  • No changes to the core agent loop struct; all hooks use the existing observer/middleware pattern
  • No new cross-subsystem coupling (e.g. src/skills/ does not import src/channels/ internals)
  • src/config/schema.rs changes are additive only; no existing keys removed or renamed
  • No personal/sensitive data in fixtures, tests, or examples; identity labels use ZeroClawAgent / zeroclaw_user only

Architecture impact

New modules

src/skills/creator.rs     # Feature A — classifier, slug sanitiser, FTS5 writer
src/skills/improver.rs    # Feature B — diff, atomic write, cooldown guard
src/tools/pipeline.rs     # Feature C — PipelineTool, step executor, interpolator

src/skills/creator.rs and src/skills/improver.rs extend the existing
src/skills/ module. They depend on the current src/skills/audit.rs
security policy and the SQLite/FTS5 backend in src/memory/ — no new
storage engines introduced.

src/tools/pipeline.rs follows the Tool naming convention and
registers via the existing tool factory in src/tools/mod.rs. It depends
only on src/tools/ and src/security/ (for the allowlist/blocklist check).

Agent loop

The core agent loop struct in src/agent/ is not modified. Features A and B
attach as post-turn observer hooks using the existing observer/middleware
pattern. Feature C is a tool registration — the loop calls it the same way it
calls any other tool.

Dependency direction

src/skills/creator.rs  →  src/skills/audit.rs   (existing, read-only)
src/skills/creator.rs  →  src/memory/            (existing FTS5 backend)
src/skills/improver.rs →  src/skills/creator.rs  (shared slug/path helpers)
src/tools/pipeline.rs  →  src/tools/             (tool registry lookup)
src/tools/pipeline.rs  →  src/security/          (allowlist/blocklist)

No imports cross the agent ↔ channels, tools ↔ providers, or
skills ↔ gateway boundaries.

Config schema

Six additive keys across two new TOML tables ([skills.auto_create],
[skills.auto_improve], [pipeline]). No existing keys are removed,
renamed, or have their types changed. Deserialisation falls back to defaults
for all new keys so existing config.toml files load without modification.

SQLite schema

One additive migration: two new columns (last_accessed_at,
last_improved_at) on the FTS5 skills table. The migration is forward-only
and does not alter existing rows or indexes.

Binary size and compile time

No new third-party crates are required. The cosine similarity check for
deduplication uses the existing embedding infrastructure already present in
src/memory/. Feature flags (skills, pipeline) gate all new code so
operators who build with --no-default-features see zero size regression.

Risk and rollback

  • Feature flags: disable skills and pipeline at compile time with --no-default-features — no runtime impact
  • Skill files are written to a new skills/ subdirectory; removing the directory fully reverts all runtime state
  • No schema migration required for existing SQLite databases — new FTS5 table is additive
  • execute_pipeline is a new tool registration; removing the factory entry reverts the feature cleanly

Breaking change?

No

Data hygiene checks

  • I removed personal/sensitive data from examples, payloads, and logs.
  • I used neutral, project-focused wording and placeholders.

Activity

  1. theonlyhennygod commented on Mar 20, 2026

    @theonlyhennygod
    Collaborator

    Thorough proposal. This is a significant feature set — effectively three separate features (auto-create, auto-improve, pipeline tool). Some thoughts:

    • Feature A (auto-create) is the most immediately useful and smallest scope. Starting here as a standalone PR would be ideal.
    • Feature B (auto-improve) has safety implications — silently patching skill files needs careful guardrails. The cooldown + similarity threshold approach is reasonable.
    • Feature C (pipeline tool) is largely independent and could be its own PR.

    Recommend splitting into 3 separate issues/PRs rather than one monolith. The skills infrastructure in src/skills/ and the tool trait in src/tools/traits.rs are the extension points. Community PRs welcome — Feature A would be a great starting point.

  2. 5queezer commented on Mar 21, 2026

    @5queezer
    ContributorAuthor

    Status update

    Feature A (autonomous skill creation) was implemented in #3916 (merged 2026-03-18) and closes #3825. The acceptance criteria for Feature A in this issue are fulfilled.

    Feature B (skill self-improvement) and Feature C (pipeline tool calling) remain unimplemented. This issue stays open to track the remaining work.

  3. changed the title [-][Feature]: Autonomous skill creation, self-improvement, and pipeline tool calling[/-] [+][Feature][learning]: Autonomous skill creation, self-improvement, and pipeline tool calling[/+] on Mar 21, 2026
  4. 5queezer commented on Mar 21, 2026

    @5queezer
    ContributorAuthor

    Feature A — criteria review (post-#3916)

    Verified PR #3916 against acceptance criteria. Checked 5/10 boxes.

    Fulfilled (checked): skill creation after qualifying turn, slug validation, cosine dedup, compile-time feature flag, tests passing.

    Remaining gaps (3):

    • created_at missing from TOML front-matter (filesystem mtime is fragile)
    • No WARN log on LRU eviction
    • No explicit path-traversal assertion (canonicalize + starts_with check)

    Should be dropped from criteria (2):

    • FTS5 indexing — contradicts this issue's own non-goal ("no cross-subsystem coupling"). Skills are discovered via filesystem scan, not SQLite.
    • cargo fuzz target for slug validation — over-engineering for a ~10-line validator. The existing unit tests with edge cases are sufficient.

    Spec deviations that are improvements over the original criteria:

    • SKILL.toml in subdirectory instead of .md (matches existing skill system convention)
    • name field instead of title (standard TOML package metadata)
    • ≥2 tool-call threshold instead of ≥3 (captures useful two-step workflows)
  5. added a commit that references this issue on Mar 24, 2026
    4aa5299
  6. added a commit that references this issue on Mar 24, 2026
    b360066
  7. added a commit that references this issue on May 24, 2026
    6f9fe1e
  8. added 2 commits that reference this issue on Oct 9, 2026
    06d8753
    0891d0c
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

enhancementNew feature or request

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions