Claude Skill MartBrowse skillsQuick linesLearn by videoTerminal guideWhat is a Skill?
← Back to list

Benchmark Checklist

Vets any performance number with seven evidence-based questions before you report it or act on it.

Dev & CodingIntermediate★ 969⑂ 112AI score 9/10Last updated: Oct 3, 2026

What it does

Whenever you produce a perf number — a PR before/after, a regression claim, a hillclimb harness, or a library/config choice — this skill makes Claude interrogate the measurement instead of trusting it.

  • Before running: write the claim in shipping words, read the measurement script to see what it times and ignores, check machine load with uptime and cores with nproc.
  • The seven questions: ① name the limiter (profile in an unreported run) ② was every side tuned like production ③ does the result break physical limits (bandwidth, cores, Amdahl arithmetic) ④ did it error, and are outputs actually correct ⑤ does it reproduce (interleaved A,B,A,B, ≥5 runs, median + range) ⑥ does it matter end-to-end ⑦ did the work even happen inside the timed region.
  • Report shape: verdict first (faster / slower / no measurable difference / inconclusive), then number with unit, run count, range, and limiter. Forces "inconclusive" when the limiter is unknown or a side ran untuned.

Who it's for

  • Backend and infra engineers who attach perf numbers to PRs
  • Teams investigating regressions or choosing between libraries and configs
  • Anyone who distrusts an agent's unverified "I optimized it" claim

Examples

  • "Swapping the JSON parser made export 30% faster" → checklist shows parsing is 1% of the request, so the real-world win is negligible.
  • Comparing two ORMs → flags that one side ran a debug build with a commit per row, so you compared configurations, not implementations.
  • A throughput number above disk bandwidth → arithmetic immediately exposes a cache or no-op, meaning the work never happened.

· · · Install guide · · ·

Try it now, no install

Paste this into Claude to use the skill without installing anything.

Read the instructions in this file and follow them to help me:
https://raw.githubusercontent.com/michael-denyer/pstack-claude/HEAD/plugins/pstack/skills/benchmark-checklist/SKILL.md

What I want: (describe your task here)

If Claude can't open the link, open it yourself and paste the contents instead.

↓ If it works for you, download the ZIP below and install it. Then it runs on its own — no pasting each time.

Install in the Claude app (no terminal)
  1. Download the ZIP with the button below.
  2. In Claude, open Settings → Capabilities and turn on 'Code execution and file creation'. (one time)
  3. Go to Customize → Skills → + → 'Upload a skill' and upload the ZIP.
↓ Download ZIP
Install in Claude Code

Let Claude do it — paste this into Claude Code

Install the skill I found on Claude Skill Mart.
Copy the plugins/pstack/skills/benchmark-checklist folder from the GitHub repo michael-denyer/pstack-claude into my ~/.claude/skills/benchmark-checklist/.
When it's done, tell me in one line what this skill can do.

Install with a command instead

git clone https://github.com/michael-denyer/pstack-claude.git /tmp/pstack-claude && mkdir -p ~/.claude/skills && cp -r /tmp/pstack-claude/plugins/pstack/skills/benchmark-checklist ~/.claude/skills/

⚠ This is a third-party skill. Check the source repository before installing.

  1. Open a terminal.
  2. Clone the repo: git clone https://github.com/michael-denyer/pstack-claude.git
  3. Create the skills folder: mkdir -p ~/.claude/skills
  4. Copy the skill: cp -r pstack-claude/plugins/pstack/skills/benchmark-checklist ~/.claude/skills/
  5. (Recommended) Copy all sibling skills so the referenced playbooks resolve: cp -r pstack-claude/plugins/pstack/skills/* ~/.claude/skills/
  6. Restart Claude Code and say something like "I just ran a benchmark, vet these results."
  7. Install perf, py-spy, strace, and pidstat beforehand so the profiling questions can actually be answered.