Skip to main content
Draft — not yet announced. This page exists so the destination and its deploy path are proven before publication day rather than on it. It carries noindex until then; remove that line in the frontmatter to publish. Figures were measured against production on 2026-08-28 and must be re-run on the day — the registry grows daily.

The finding

We publish a lift number for agent skills: the pass rate with the skill installed, minus the pass rate without it, on the same case set. Forty-eight public skills in our registry carry that number on two different models. We compared them. All forty-eight moved. Not one held its number. Thirty-six fell, twelve rose, mean change −17.3 points. The tempting headline is “lift decays with newer models.” We are not making that claim, because our own data refutes it — a quarter of these skills got better. The honest finding is narrower, and more useful:
A skill’s lift is not a property of the skill. It is a property of the skill, the model, the harness and the case set, together. Report it without pinning those and you have published a number nobody can reproduce — including you, three weeks later.

The cleanest example

The obvious objection is that two runs might not be comparable. It is a fair objection, and it applies to most of our pairs — see Limits. So here is the finding on a pair where it cannot apply. postgresql-conventions ran the same 24 cases on both models: Same skill, same suite, same grader. One model upgrade, and four fifths of the measured benefit is gone. Nothing about the skill changed, and nothing about it was wrong. The number was not portable, and nothing on its scorecard said so. Two other pairs share a case set and behave the same way: svg-ui-panel-conventions (91.3 → 73.9 on 23 cases) and thesis-dissertation-compliance (50.0 → 41.7 on 24 cases).

Method

The corpus is ours, and it is synthetic. These runs were produced by our own evaluation fleet under pinned and identical conditions — not by third-party production traffic. That makes it a measuring instrument, which is what this report uses it as. It would not support a claim about adoption, and we are not making one. Each pair is the most recent sound benchmark run per (skill, model). A run is sound when its suite survived calibration and produced a headline over at least 8 cases.

Limits, stated before you find them

  • 45 of the 48 pairs ran different case counts on the two models, because suites get re-cut over time. That is exactly why the worked example above is one of the three that did not. Case-set drift moves a pass rate on its own, so treat the population figure as directional and the three clean pairs as the evidence.
  • Two models, one family. We have not shown this holds across vendors.
  • Our headline floor is 8 cases. An 8-case suite can move 12.5 points on one case. Small-n pairs are in the population figure and are not in the worked example.
  • We publish these scores ourselves. That is a conflict, and stating it is cheaper than having it pointed out. Everything below ships so you can recompute it.

What ships with this

  • The raw per-case rows for all 48 pairs — not a summary table.
  • An agentversion manifest pinning model, harness and case set for every row. That is the whole point: it is what makes a lift number comparable, and it is why these pairs could be compared at all.
  • A one-command reproduction.