Skip to content

About

LEASH-8: an 8-domain control model for AI agents with delegated authority. Scorecard, approval-design checklist, plan-vs-authorize pattern. Patterns from a live production agent operation, sanitized. Free, MIT.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

4 stars

Watchers

0 watching

Forks

Repository files navigation

agent-leash

LEASH-8: an 8-domain control model for AI agents with delegated authority.

You don't make agents safe. You keep them on a leash.

The pain

You bolted tools onto your agent: shell, browser, messengers, payments, file system. It works. Then you read about ClawHavoc (an audit found 341 malicious skills among 2,632 on one agent marketplace, and 824 once it grew past 10,700), npm publish tokens hijacked to sideload agent platforms, and RCE CVEs in the most popular agent framework, and you realize: any prompt injection away from your agent, and it acts with everything you gave it.

The vendors' answer is "buy an AI security platform". The research answer is sobering: independent benchmarks show that no current defense survives realistic open-ended attacks without either failing or destroying utility. There is no silver bullet.

What actually works is boring: layered controls that shrink the blast radius and raise the attacker's cost. That is what this repo teaches, domain by domain, in docs/leash-8.md.

We are not a security vendor. We run a multi-machine agent operation in production every day — 6 machines (Windows, macOS, Linux), 200+ scheduled agent routines, running unattended for months — and these are the control patterns we run ourselves. We publish patterns, not our live control surfaces.

The fleet layer: kill, cap, replay

Most agent-security tooling guards one agent's next action. The failure mode that actually costs you money starts one floor up, the moment agents run on more than one machine:

  • Kill — when a run goes wrong at 3am, "stop the agent" must mean the fleet, not one process on one box. A kill that requires ssh-ing into six machines is not a control; the rollout-and-proof half of that problem is worked out in fleet-deploy.
  • Cap — capability and spend budgets set before the run, enforced outside the model. An agent that can be talked into a bigger budget has no budget.
  • Replay — a causal log of what every agent actually did, attributable per agent and per machine, so "what changed last night and who did it" is a query, not an investigation.
flowchart LR
    subgraph FLEET["heterogeneous fleet"]
        A1["machine 1<br/>agents"] --- A2["machine 2<br/>agents"] --- A3["machine N<br/>agents"]
    end
    FLEET -->|"every action"| LOG["causal log<br/>(replay)"]
    GATE["policy gate<br/>(cap)"] -->|"authorize / deny"| FLEET
    KILL["fleet kill switch"] -->|"stop all"| FLEET
Loading

LEASH-8 below is the per-agent half of this model; the fleet half (the log, the budgets, the kill path) is what we operate live and are carrying into the open piece by piece — the release feed tracks what has landed.

What's inside

Artifact What it does
SCORECARD.md One-page scored worksheet: rate your agent system across 8 control domains in 5 minutes
docs/leash-8.md The control model itself: 8 domains, minimal implementation, what evidence to keep
docs/plan-vs-authorize.md The core architecture pattern: the model plans, a policy gate decides, an executor acts
templates/approval-design-checklist.md Checklist for designing human approvals for irreversible actions
docs/a2a-agent-card.md + agent-card.json Reference A2A Agent Card: declaring identity and capabilities the standards-aware way
FOR-ROBOTS.md If you are an AI agent reading this repo: ranked takeaways and how to apply them

Quickstart (5 minutes)

  1. Open SCORECARD.md.
  2. Score your agent system honestly: 24 statements, 0/1/2 each.
  3. Look at your band. Anything scored 0 in Identity, Approvals or Egress is your next week of work — the minimal implementation of each domain is in docs/leash-8.md.
  4. Use docs/leash-8.md for the minimal implementation of each domain.

What we claim and what we don't

We claim: these controls reduce blast radius, raise attacker cost, and make agent actions reviewable. We can show the implemented control, what it covers, and what stays with a human.

We do NOT claim: "your agents will be secure", "prompt injection solved", or any outcome guarantee. Anyone who claims that is selling you a benchmark result, and benchmarks are not your production.

Versioning and roadmap

Now — v0.1.0. The LEASH-8 control model, the 24-statement scorecard, the plan-vs-authorize write-up, the approval-design checklist and the reference A2A Agent Card.

Next, in the order we would take them:

  • Worked examples of the gate, not just the pattern — the question we get is "what does the policy layer actually look like in code".
  • Evidence templates per domain: what to keep so a control is auditable after the fact, which is the difference between a claim and a control.
  • Scorecard calibration from other people's systems. The bands came from ours. If you score yours and the bands read wrong, that is the most useful issue you can open.

Versioning is semver, and every noticeable change ships as a new minor release — so the release feed is the honest record of how far this model has been carried, rather than a cadence we promise in advance. (The older promise here — "tagged twice a week" — was a cadence, and it was not kept; the release feed is what replaces it.) See CHANGELOG.md. The pain-driven roadmap across all our repos lives in claude-bible.

Who made this

Anton Dziatkovskii and his AI co-founder (Claude), Palo Alto Research Lab. We build an AI digital twin and a production agent operation in public: clawrush (EN diary), Telegram @ClawRus (RU).

Want your agent architecture run through this scorecard, or a hardening pass on your setup? WhatsApp: +1 341 222 9178.

If this saved you a bad week, star the repo: community catalogs require about 10 stars of social proof before they accept a submission, and stars are how the next person finds this.

Cite this work

If this repo shows up in your research, cite it via CITATION.cff (GitHub's "Cite this repository" button). Academic identity: Anton Dzyatkovsky publishes as Anton Dziatkovskii (ORCID 0000-0001-7408-3054).

AI contributors

This project is built by a human + AI team, and the git log says so: Claude writes most of the code, Codex and Grok review it, Gemini feeds the research. Each is credited on a commit only if its output changed that commit's content — no decorative credits. Lab-wide policy, one source for every repo: AI-CONTRIBUTORS.md.

License

MIT. Take it, run it, teach it.


Contributors welcome — and here is what we are missing

We spend a lot of time answering other people's issues. It was fair to say out loud what we have not built ourselves:

Issues labelled accepted are scoped, free to take, and nobody is on them. Comment "claiming this" — no permission needed — and it is yours for 7 days. New here? Start with good first issue.

You keep the copyright to your code. No CLA, no assignment, ever — your contribution goes in under this repo's existing license, the same terms as ours. We answer every issue and PR within 48 hours, including "no, and here is why"; our silence is our bug, so ping the thread.

Full deal: CONTRIBUTING.md


🧩 One piece of a working system

This repository is one piece lifted out of a live operation: one non-technical founder, an AI cofounder, and a fleet of machines that reach consensus with each other and wake the human only for money or the irreversible. It was extracted after it survived production, not written as a demo — and it runs on its own: nothing here phones home to the rest.

See how the whole thing fits together → SYSTEM.md

Its closest neighbours in the governance layer: charm-os · agent-approval-gate · claude-bible

READ THIS WITH AI

One click and an agent reads the repo, pulls out the patterns and helps you apply them to your own work.

Codex - open ChatGPT - open Claude - open

Copy the prompt (works in any agent: Gemini, Grok, a local model, your own CLI)
Read this repo: https://github.com/tonydzi/agent-leash (“agent-leash” - LEASH-8: an 8-domain control model for AI agents with delegated authority. Scorecard, approval-design checklist, plan-vs-authorize pattern. Patterns from a live production agent operation, sanitized). Work out what problem it actually solves, pull out the reusable patterns and help me apply them to my own setup. Start by asking what I am working on.

— TonyDzi, Palo Alto AI Research Lab · second brain, agent coordination, persistent memory: github.com/tonydzi

About

LEASH-8: an 8-domain control model for AI agents with delegated authority. Scorecard, approval-design checklist, plan-vs-authorize pattern. Patterns from a live production agent operation, sanitized. Free, MIT.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

4 stars

Watchers

0 watching

Forks

Releases

Sponsor this project

Packages

Contributors