Benchmark Checklist
Vets any performance number with seven evidence-based questions before you report it or act on it.
Dev & CodingIntermediate★ 969⑂ 112AI score 9/10Last updated: Oct 3, 2026
What it does
Whenever you produce a perf number — a PR before/after, a regression claim, a hillclimb harness, or a library/config choice — this skill makes Claude interrogate the measurement instead of trusting it.
- Before running: write the claim in shipping words, read the measurement script to see what it times and ignores, check machine load with
uptimeand cores withnproc. - The seven questions: ① name the limiter (profile in an unreported run) ② was every side tuned like production ③ does the result break physical limits (bandwidth, cores, Amdahl arithmetic) ④ did it error, and are outputs actually correct ⑤ does it reproduce (interleaved A,B,A,B, ≥5 runs, median + range) ⑥ does it matter end-to-end ⑦ did the work even happen inside the timed region.
- Report shape: verdict first (faster / slower / no measurable difference / inconclusive), then number with unit, run count, range, and limiter. Forces "inconclusive" when the limiter is unknown or a side ran untuned.
Who it's for
- Backend and infra engineers who attach perf numbers to PRs
- Teams investigating regressions or choosing between libraries and configs
- Anyone who distrusts an agent's unverified "I optimized it" claim
Examples
- "Swapping the JSON parser made export 30% faster" → checklist shows parsing is 1% of the request, so the real-world win is negligible.
- Comparing two ORMs → flags that one side ran a debug build with a commit per row, so you compared configurations, not implementations.
- A throughput number above disk bandwidth → arithmetic immediately exposes a cache or no-op, meaning the work never happened.
· · · Install guide · · ·
Try it now, no install
Paste this into Claude to use the skill without installing anything.
Read the instructions in this file and follow them to help me: https://raw.githubusercontent.com/michael-denyer/pstack-claude/HEAD/plugins/pstack/skills/benchmark-checklist/SKILL.md What I want: (describe your task here)
If Claude can't open the link, open it yourself and paste the contents instead.
↓ If it works for you, download the ZIP below and install it. Then it runs on its own — no pasting each time.
Install in the Claude app (no terminal)
- Download the ZIP with the button below.
- In Claude, open Settings → Capabilities and turn on 'Code execution and file creation'. (one time)
- Go to Customize → Skills → + → 'Upload a skill' and upload the ZIP.
Install in Claude Code
Let Claude do it — paste this into Claude Code
Install the skill I found on Claude Skill Mart. Copy the plugins/pstack/skills/benchmark-checklist folder from the GitHub repo michael-denyer/pstack-claude into my ~/.claude/skills/benchmark-checklist/. When it's done, tell me in one line what this skill can do.
Install with a command instead
git clone https://github.com/michael-denyer/pstack-claude.git /tmp/pstack-claude && mkdir -p ~/.claude/skills && cp -r /tmp/pstack-claude/plugins/pstack/skills/benchmark-checklist ~/.claude/skills/⚠ This is a third-party skill. Check the source repository before installing.
- Open a terminal.
- Clone the repo:
git clone https://github.com/michael-denyer/pstack-claude.git - Create the skills folder:
mkdir -p ~/.claude/skills - Copy the skill:
cp -r pstack-claude/plugins/pstack/skills/benchmark-checklist ~/.claude/skills/ - (Recommended) Copy all sibling skills so the referenced playbooks resolve:
cp -r pstack-claude/plugins/pstack/skills/* ~/.claude/skills/ - Restart Claude Code and say something like "I just ran a benchmark, vet these results."
- Install
perf,py-spy,strace, andpidstatbeforehand so the profiling questions can actually be answered.
View source on GitHub ↗License: MIT