Agent Benchmarking
Head-to-head coding agent comparison tool: YAML task definitions with judge criteria, git worktree isolation per agent run, pass rate/cost/time/consistency metrics, and reproducible benchmarking across Claude Code, Aider, Codex, and other agents.
Governance Receipt
- Signer
- sovereign-claw-ed25519
- Signed At
- 6/4/2026
- Risk Tier
- T2
- Receipt Hash
- cee94f86
- Manifest Hash
- 9bb29a8b037ffd9dd19de60289d7292858c4c5cc8688632eaa09e9be2e9c7d25
- Signature
- s/aDylbU
- Root Public Key
- 349b0348
Skill Details
- Gate Verdict
- Elevated · Review
- Publication State
- published
- Risk Tier
- T2
- Manifest Hash
- 9bb29a8b
Download Attested Bundle
Install Agent Benchmarking
Each bundle is served from the attested copy of Agent Benchmarking, so the file you download is the one covered by the governance receipt on this page rather than whatever currently sits at the upstream source.
Claude Code skill
curl -L "https://atestiv.com/api/download/ecc-agent-benchmarking?format=claude-code" -o ecc-agent-benchmarking.skill.mdMCP tool definitions
curl -L "https://atestiv.com/api/download/ecc-agent-benchmarking?format=mcp" -o ecc-agent-benchmarking.mcp-tools.jsonRaw attested bundle
curl -L "https://atestiv.com/api/download/ecc-agent-benchmarking?format=raw" -o ecc-agent-benchmarking.bundle.jsonVerify the bundle against its manifest hash
shasum -a 256 ecc-agent-benchmarking.bundle.jsonProvenance
- Catalog identifier:
ecc-agent-benchmarking - Harvested from the upstream repository.
- The manifest hash below is what the signing key committed to. A bundle whose digest does not match it is not the attested Agent Benchmarking.
- Anything published here has been through the same gate, so a missing receipt is itself a signal rather than an omission.
More Skills
Eval-Driven Development
Eval-driven development framework: define pass/fail criteria before implementation, capability/regression/consistency eval types, pass@k reliability metrics, grader patterns, and continuous eval integration during development.
agent-designer
Use when the user asks to design multi-agent systems, create agent architectures, define agent communication patterns, or build autonomous agent workflows.
Agent Lifecycle Ops
Operate long-lived agent workloads with observability, security boundaries, and lifecycle management patterns for enterprise deployments.