Claude Skill MartBrowse skillsQuick linesLearn by videoTerminal guideWhat is a Skill?
← Back to list

Evals Bootstrap

Turns an agent's real failures into a fast, trace-based behavioural eval suite with a one-command runner.

Dev & CodingIntermediate★ 815⑂ 128AI score 9/10Last updated: Sep 29, 2026

What it does

  • Mines your agent's logs or transcripts for up to twenty real failures and clusters them into four to eight named behaviours.
  • Writes cases.yaml, one entry per failure: input, a human-readable expect, and machine-checked check rules (trace has X, trace has X before Y, trace lacks X).
  • Sets up trace storage (traces/<id>.json) if runs aren't logged yet — no trace, no behavioural checks.
  • Generates a dependency-light check_traces.py runner (pyyaml only, no model calls) that prints per-case results and exits non-zero on failure, so it can gate CI.
  • Hands back the flywheel: read fresh traces weekly, add each new failure as a case, fix the biggest cluster, re-run.

Who it's for

  • Engineers shipping tool-calling LLM agents who fear regressions after prompt or model swaps.
  • Teams with no evaluation harness yet who need a first golden set.
  • Anyone who prefers deterministic assertions over LLM-as-judge scoring.

Examples

  1. "I changed the system prompt — did anything break?" → the skill mines past transcripts and builds a regression suite.
  2. A refund agent replied without looking up the order → add trace has get_order before send_reply.
  3. A refund executed without approval → lock it with trace has approval_request before refund and wire python check_traces.py into CI.

· · · Install guide · · ·

Try it now, no install

Paste this into Claude to use the skill without installing anything.

Read the instructions in this file and follow them to help me:
https://raw.githubusercontent.com/undefined-ui/second-brain-os/HEAD/plugins/agents-course/skills/evals-bootstrap/SKILL.md

What I want: (describe your task here)

If Claude can't open the link, open it yourself and paste the contents instead.

↓ If it works for you, download the ZIP below and install it. Then it runs on its own — no pasting each time.

Install in the Claude app (no terminal)
  1. Download the ZIP with the button below.
  2. In Claude, open Settings → Capabilities and turn on 'Code execution and file creation'. (one time)
  3. Go to Customize → Skills → + → 'Upload a skill' and upload the ZIP.
↓ Download ZIP
Install in Claude Code

Let Claude do it — paste this into Claude Code

Install the skill I found on Claude Skill Mart.
Copy the plugins/agents-course/skills/evals-bootstrap folder from the GitHub repo undefined-ui/second-brain-os into my ~/.claude/skills/evals-bootstrap/.
When it's done, tell me in one line what this skill can do.

Install with a command instead

git clone https://github.com/undefined-ui/second-brain-os.git && mkdir -p ~/.claude/skills && cp -r second-brain-os/plugins/agents-course/skills/evals-bootstrap ~/.claude/skills/

⚠ This is a third-party skill. Check the source repository before installing.

  1. Open a terminal.
  2. Clone the repo: git clone https://github.com/undefined-ui/second-brain-os.git
  3. Create the skills directory: mkdir -p ~/.claude/skills
  4. Copy the skill: cp -r second-brain-os/plugins/agents-course/skills/evals-bootstrap ~/.claude/skills/
  5. Make sure Python 3 is available and install the one dependency: pip install pyyaml
  6. Restart Claude Code, then ask something like "bootstrap an eval suite for my agent".
  7. Point it at where your runs are logged (transcripts or a traces directory) so it can extract real cases.