Skip to content

About

Safety-first MCP server and portable agent skillbox for authorized local AI red-team evaluation.

Resources

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Repository files navigation

redteam-skillbox logo

redteam-skillbox

CI

redteam-skillbox is a Python 3.11+ MCP server and portable agent skillbox for authorized, local, disposable AI red-team evaluation.

It is designed for agent-capable models across broad provider infrastructures and local model agents. The core contract is capability-based: if an agent can read local skill files, call MCP tools or equivalent strict JSON Schema tools, run approved local commands, and preserve sandbox evidence, it can use the skillbox. Product-specific manifests are optional compatibility adapters, not the source of truth.

It ships:

  • Three portable agent skills under skills/.
  • A validator CLI: redteam-skillbox validate.
  • A local stdio MCP server entrypoint: redteam-skillbox-mcp --root PATH.
  • A read-only local document connector with search and fetch.
  • A general MCP runtime with classified tools.
  • JSON Schemas for agent manifests, eval cases, skill-selection cases, evidence reports, sandbox manifests, and MCP tool registry metadata.

Quickstart

py -3.11 -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install -e ".[dev]"
ruff check .
pytest
redteam-skillbox validate

On Unix-like shells:

python3.11 -m venv .venv
. .venv/bin/activate
python -m pip install -e ".[dev]"
ruff check .
pytest
redteam-skillbox validate

Validate an adapted local project skill pack:

redteam-skillbox validate --root /path/to/project
redteam-skillbox-mcp --root /path/to/project

Safety Boundary

The project is only for authorized, local, disposable evaluation. It refuses or redirects requests involving real credentials, real customer projects, production systems, third-party targets, public package indexes as attack surfaces, shell startup files, Git hooks, global language environments, persistent host indexes, external service probing, credential replay, exploit payload generation, malware, persistence, evasion, or exfiltration.

Safe redirect wording:

I can't run or design that against real systems or credentials. I can help convert it into an authorized sandbox evaluation using generated fixtures, loopback-only services, and a disposable sandbox root.

MCP Modes

  • Read-only connector: search(query, filters?) and fetch(id) over local project docs, skills, examples, and non-sensitive reports. Exposed as in-process helpers (LocalDocumentIndex, redteam_skillbox.mcp.resources) and as MCP tools on the general server.
  • General server: classified tools for skill listing, validation, sandbox creation, generated local eval execution, evidence grading, report rendering, and manifest-verified cleanup.

network_posture and approval_required are advisory metadata honored by the agent, MCP client, and host, not server-enforced controls. For mechanical loopback-only isolation, run the server inside an OS network sandbox (see docs/deployment.md). Path containment, manifest-verified cleanup, generated-fixture/loopback gating, and the scrubbed fixture environment are enforced in code. See docs/threat-model.md.

Agent Compatibility

Each skill includes agents/agent.yaml, a vendor-neutral manifest with display metadata, required capabilities, compatible runtime classes, invocation modes, and safety policy. agents/openai.yaml is included only as an optional compatibility adapter for clients that understand that file.

Target runtimes include local MCP-capable agents, hosted agent tool runtimes, filesystem-reading agents, filesystem-writing sandbox agents, and tool-calling model agents. The project does not assume a specific model provider.

See docs/agent-compatibility.md for the runtime contract. See docs/adapter-architecture.md for projection rules across MCP, hosted connector, native tool, and OpenAPI-style runtimes. See docs/skill-authoring.md for adapting the portable skills inside a local project. See docs/deployment.md for PyPI, uvx, container, and MCP client configuration. See docs/redis-sidecar.md for optional sandbox-scoped Redis context, process appendage, and report delivery. Redis remains disabled by default and requires redteam-skillbox[redis] plus a local redis-server binary when enabled.

Public Release

This is a v1 public beta while runtime isolation and adapter boundaries continue to harden.

Release notes are in RELEASE.md.

Evidence

Every stateful eval run must produce an evidence report with run ID, timestamp, version/ref, skill invoked, explicit invocation text, sandbox root, manifest path, network posture, commands or tool calls, cwd, environment overrides, exit codes, stdout/stderr excerpts, observations, findings, unsupported claims, cleanup status, provenance metadata, and runtime-context metadata.

About

Safety-first MCP server and portable agent skillbox for authorized local AI red-team evaluation.

Resources

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages