Evals Bootstrap
Turns an agent's real failures into a fast, trace-based behavioural eval suite with a one-command runner.
Dev & CodingIntermediate★ 815⑂ 128AI score 9/10Last updated: Sep 29, 2026
What it does
- Mines your agent's logs or transcripts for up to twenty real failures and clusters them into four to eight named behaviours.
- Writes
cases.yaml, one entry per failure:input, a human-readableexpect, and machine-checkedcheckrules (trace has X,trace has X before Y,trace lacks X). - Sets up trace storage (
traces/<id>.json) if runs aren't logged yet — no trace, no behavioural checks. - Generates a dependency-light
check_traces.pyrunner (pyyaml only, no model calls) that prints per-case results and exits non-zero on failure, so it can gate CI. - Hands back the flywheel: read fresh traces weekly, add each new failure as a case, fix the biggest cluster, re-run.
Who it's for
- Engineers shipping tool-calling LLM agents who fear regressions after prompt or model swaps.
- Teams with no evaluation harness yet who need a first golden set.
- Anyone who prefers deterministic assertions over LLM-as-judge scoring.
Examples
- "I changed the system prompt — did anything break?" → the skill mines past transcripts and builds a regression suite.
- A refund agent replied without looking up the order → add
trace has get_order before send_reply. - A refund executed without approval → lock it with
trace has approval_request before refundand wirepython check_traces.pyinto CI.
· · · Install guide · · ·
Try it now, no install
Paste this into Claude to use the skill without installing anything.
Read the instructions in this file and follow them to help me: https://raw.githubusercontent.com/undefined-ui/second-brain-os/HEAD/plugins/agents-course/skills/evals-bootstrap/SKILL.md What I want: (describe your task here)
If Claude can't open the link, open it yourself and paste the contents instead.
↓ If it works for you, download the ZIP below and install it. Then it runs on its own — no pasting each time.
Install in the Claude app (no terminal)
- Download the ZIP with the button below.
- In Claude, open Settings → Capabilities and turn on 'Code execution and file creation'. (one time)
- Go to Customize → Skills → + → 'Upload a skill' and upload the ZIP.
Install in Claude Code
Let Claude do it — paste this into Claude Code
Install the skill I found on Claude Skill Mart. Copy the plugins/agents-course/skills/evals-bootstrap folder from the GitHub repo undefined-ui/second-brain-os into my ~/.claude/skills/evals-bootstrap/. When it's done, tell me in one line what this skill can do.
Install with a command instead
git clone https://github.com/undefined-ui/second-brain-os.git && mkdir -p ~/.claude/skills && cp -r second-brain-os/plugins/agents-course/skills/evals-bootstrap ~/.claude/skills/⚠ This is a third-party skill. Check the source repository before installing.
- Open a terminal.
- Clone the repo:
git clone https://github.com/undefined-ui/second-brain-os.git - Create the skills directory:
mkdir -p ~/.claude/skills - Copy the skill:
cp -r second-brain-os/plugins/agents-course/skills/evals-bootstrap ~/.claude/skills/ - Make sure Python 3 is available and install the one dependency:
pip install pyyaml - Restart Claude Code, then ask something like "bootstrap an eval suite for my agent".
- Point it at where your runs are logged (transcripts or a traces directory) so it can extract real cases.
View source on GitHub ↗License: MIT