Back to feed
Dev.to
Dev.to
7/24/2026
I benchmarked Claude Code skills against a placebo — and half of mine failed

I benchmarked Claude Code skills against a placebo — and half of mine failed

Short summary

The author benchmarked Claude Code agent skills using a rigorous three-arm methodology (off/placebo/on) across 516 runs on Claude Opus 4. Half the tested skills failed to beat a placebo prompt, including popular ones. A focused 20-line anti-over-engineering skill cut source LOC by 23.8% at identical accuracy, outperforming the 196k-star Karpathy Guidelines on that specific axis. The key insight: instruction mechanism matters, not just instruction presence.

  • 516-run benchmark with placebo-controlled methodology for Claude Code agent skills
  • Half of tested skills failed to beat a same-length placebo prompt
  • Focused 20-line skill outperformed 196k-star Karpathy Guidelines on code minimality at equal accuracy

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more