Behaviour packs for Claude Code
Does the code it writes actually build?
1 of 5 packs separates from installing nothing on frontend work; 3 cannot be told apart from it at these sample sizes, and 1 wrote too little code to grade. That is the result, not a missing one — every pack here was installed and run against real tickets in a real repository, unattended, and the code it delivered was put through that repository's own build. There is no submission form and no self-reported score.
Frontend build rate · blocked condition
Share of tickets whose delivered code built
Ordered by point estimate, with the Wilson interval drawn on every bar. The overlap is the point: read the order, then read how much of it survives a corpus this size.
Clears the baseline — its interval sits entirely above the baseline's
Not distinguishable from the baseline at this sample size — intervals overlap
Not measured on frontend work — no graded cell, so no interval
Wilson interval bar = point estimate dashed track = baseline (no pack installed) tier = whether this corpus can tell the arm apart from the baseline
What these numbers are over. 12 tasks, all in one repository — tiangolo/full-stack-fastapi-template at cd83fc1. 6 tmpl-be (102 cells), 6 tmpl-fe (127 cells). Every frontend task adds a React component to that template; every backend task adds an endpoint to it, or a constraint on one. There is no mobile or client-app task here, no design task, no refactor, no debugging, no test authoring, and no task in a second repository. This page used to say less than that.
So these results describe this kind of work in this repository. Whether they transfer to other repositories, other stacks, or other kinds of ticket is not something anything here has measured — and a benchmark that lets you assume it did is doing the thing this one exists to object to.
If you came to decide. On frontend work in this corpus, ponytail is the only arm whose interval clears the baseline's outright — every other pack overlaps it, so at these sample sizes they are not distinguishable from installing nothing. It is not the cheapest way to get there: it costs more per cell than the baseline, so the choice is a trade rather than a free win. On backend work, no pack separates from the baseline at all — every interval overlaps it. Nothing on this page supports choosing a pack for backend tickets.
There is no composite index here, and that is deliberate. A single score would need weights across build rate, cost, turns and silence, and this corpus gives no basis for choosing them — an index would be our opinion wearing a number's clothes. What the ordering above supports is one sentence: of the measured arms, ponytail clears the baseline's interval outright. Every other interval overlaps the baseline's, so at these sample sizes those packs are not distinguishable from installing nothing.
Cost against build rate
What a build is worth paying for
Mean cost per cell against the frontend build rate. Up and to the left is better; the shaded region is the quadrant that beats the baseline on both.
Build rate against mean cost per cell
Higher and further left is better.
The dashed line is the Pareto frontier: an arm below and to the right of it is beaten on both axes at once by something already on this chart.
- baseline
- caveman
- karpathy
- mattpocock
- ponytail
- superpowers
The shaded quadrant is empty, and that is the result. No pack here is both cheaper per cell than installing nothing and better at getting code to build. The one arm that clears the baseline on build rate costs more per cell to do it, so the choice is a trade rather than a free win — which is why this page draws the axes instead of ranking them into one number.
Grouped by what this corpus can separate
The packs
Each card states what this corpus found for that pack, on each kind of work, with the counts it rests on. Cards are grouped by whether the corpus can tell the pack apart from installing nothing, and stay alphabetical inside a group — what cannot be separated cannot be ordered.
5 installing, 0 blocked — measured by a real install into a scratch config directory, oldest attempt 2026-08-18. A listing that stopped installing keeps its card and its figures — the measurement happened; what changed is upstream.
Separates from installing nothing 1
Its interval sits entirely above the no-pack interval on frontend work.
-
ponytail
Built 16 of 22 frontend tickets, against 5 of 20 with no pack installed. Its interval sits entirely above the no-pack interval, so this corpus does separate them.
Passed 8 of 17 backend tickets, against 7 of 15 with no pack installed. The two intervals overlap, so this corpus cannot tell them apart at these sample sizes.
No pack here separates from no pack on backend work. That gate asks whether the code imports and adds no new diagnostic; it runs no test suite.
73% n=22 (frontend), 95% interval 52% - 87% no pack 25% n=20 (frontend) 47% n=17 (backend), 95% interval 26% - 69% no pack 47% n=15 (backend)
installs
No difference this corpus can measure 3
Every interval here overlaps the no-pack interval, so these are not ranked against one another — this corpus cannot order what it cannot separate. That is a statement about the sample, not a finding about any of these packs.
-
andrej-karpathy-skills
Built 8 of 20 frontend tickets, against 5 of 20 with no pack installed. The two intervals overlap, so this corpus cannot tell them apart at these sample sizes.
Passed 7 of 18 backend tickets, against 7 of 15 with no pack installed. The two intervals overlap, so this corpus cannot tell them apart at these sample sizes.
No pack here separates from no pack on backend work. That gate asks whether the code imports and adds no new diagnostic; it runs no test suite.
40% n=20 (frontend), 95% interval 22% - 61% no pack 25% n=20 (frontend) 39% n=18 (backend), 95% interval 20% - 61% no pack 47% n=15 (backend)
installs
-
caveman
Built 6 of 21 frontend tickets, against 5 of 20 with no pack installed. The two intervals overlap, so this corpus cannot tell them apart at these sample sizes.
Passed 6 of 17 backend tickets, against 7 of 15 with no pack installed. The two intervals overlap, so this corpus cannot tell them apart at these sample sizes.
No pack here separates from no pack on backend work. That gate asks whether the code imports and adds no new diagnostic; it runs no test suite.
29% n=21 (frontend), 95% interval 14% - 50% no pack 25% n=20 (frontend) 35% n=17 (backend), 95% interval 17% - 59% no pack 47% n=15 (backend)
installs
-
mattpocock-skills
Built 10 of 20 frontend tickets, against 5 of 20 with no pack installed. The two intervals overlap, so this corpus cannot tell them apart at these sample sizes.
Passed 7 of 18 backend tickets, against 7 of 15 with no pack installed. The two intervals overlap, so this corpus cannot tell them apart at these sample sizes.
No pack here separates from no pack on backend work. That gate asks whether the code imports and adds no new diagnostic; it runs no test suite.
50% n=20 (frontend), 95% interval 30% - 70% no pack 25% n=20 (frontend) 39% n=18 (backend), 95% interval 20% - 61% no pack 47% n=15 (backend)
installs
Nothing to grade 1
Too few graded cells for any difference from no pack to have shown up at all.
-
superpowers
No frontend tickets were graded for this pack, so there is no rate to compare. What the runs did produce is stated below.
Nothing to grade on backend tickets. This corpus graded 0 of 1, too few for any difference from no pack to have shown up at all — so no verdict is offered rather than a weak one.
writes no code in 98% of cells, n=41installs
Method
What these numbers are, exactly
Every figure above is the blocked condition: the agent was given the ticket with Bash withheld and told to write rather than run. A second condition was measured with Bash allowed, at a smaller sample and without one of the arms. The two are not averaged together — the gap between their baselines is as wide as the widest gap any pack opens on its own baseline — so this page reports one condition and names it rather than publishing a rate that describes neither.
The headline is the frontend build rate: the share of frontend tickets whose delivered code the repository's own build accepted. The backend tickets are graded by a different gate — whether the delivered code imports and introduces no new lint or type diagnostic — so they are printed beside it rather than folded into it. Two populations, two denominators, no single number spanning both. The analysis page shows what each of them is made of, cell by cell.
Runs withdrawn as instrument failures are excluded and stay on record; the count of them is 1.
Withdrawn: what this page used to say about scope
Until 2026-08-20 this page said only that every pack was run against real tickets in a real repository, and stopped there. That sentence is true and it is read as far wider than what happened — a reader pictures a range of work, and the corpus is one template repository and two shapes of task. Nothing was measured that the page did not report; what the page did not report was the shape of what it measured, which is the same defect as a rate without its denominator.
The cause: nobody had counted. The scope was discoverable from the analysis page and from the records, and never stated beside the headline where a reader forms an impression. It was found by the project's owner and, independently, by an outside review of this site. The replacement is the paragraph beside the chart above, derived from the records so a corpus change rewrites it rather than outdating it.