Agent Benchmarking

Elevated · Review

Head-to-head coding agent comparison tool: YAML task definitions with judge criteria, git worktree isolation per agent run, pass rate/cost/time/consistency metrics, and reproducible benchmarking across Claude Code, Aider, Codex, and other agents.

Governance Receipt

Signer
sovereign-claw-ed25519
Signed At
6/4/2026
Risk Tier
T2
Receipt Hash
cee94f86
Manifest Hash
9bb29a8b037ffd9dd19de60289d7292858c4c5cc8688632eaa09e9be2e9c7d25
Signature
s/aDylbU
Root Public Key
349b0348

Skill Details

Gate Verdict
Elevated · Review
Publication State
published
Risk Tier
T2
Manifest Hash
9bb29a8b

Install Agent Benchmarking

Each bundle is served from the attested copy of Agent Benchmarking, so the file you download is the one covered by the governance receipt on this page rather than whatever currently sits at the upstream source.

Claude Code skill

curl -L "https://atestiv.com/api/download/ecc-agent-benchmarking?format=claude-code" -o ecc-agent-benchmarking.skill.md

MCP tool definitions

curl -L "https://atestiv.com/api/download/ecc-agent-benchmarking?format=mcp" -o ecc-agent-benchmarking.mcp-tools.json

Raw attested bundle

curl -L "https://atestiv.com/api/download/ecc-agent-benchmarking?format=raw" -o ecc-agent-benchmarking.bundle.json

Verify the bundle against its manifest hash

shasum -a 256 ecc-agent-benchmarking.bundle.json

Provenance

  • Catalog identifier: ecc-agent-benchmarking
  • Harvested from the upstream repository.
  • The manifest hash below is what the signing key committed to. A bundle whose digest does not match it is not the attested Agent Benchmarking.
  • Anything published here has been through the same gate, so a missing receipt is itself a signal rather than an omission.

More Skills