1 August 2026
We expected a gap between cheap and frontier models. There isn’t one
Synthetic narration · 6 min
unPaper’s template carries a catalogue of layouts: a waterfall, a coordinate plot, a network map, a span bar, a ring. We built each one because models kept constructing it by hand. The obvious follow-up question is whether anyone actually uses them — so we counted, per model, how often a deck used a named layout when the content called for one.
What we found
Every model we tested — the cheapest open-weight ones and the two most expensive frontier ones — reaches for a named layout in four or five briefs out of five. The two frontier models sit at the top of the table, not the bottom.
- claude-opus-5primed — see method100% (5 decks)
- gpt-5.2100% (5 decks)
- gpt-oss-120b100% (5 decks)
- mistral-small-3.2-24b-instruct100% (5 decks)
- deepseek-v4-flash80% (5 decks)
- glm-5.280% (5 decks)
The gap we expected, and went looking for
This entry was going to be about a difference. We believed, and had written down, that cheap models use the catalogue about twice as often as frontier ones — and that this was unPaper working exactly as designed, because the instructions invite a capable model to compose beyond the catalogue and a weaker one to reach into it.
Checking that before publishing killed it. Every frontier deck we were counting had been built before most of these layouts existed, so the models were being marked down for failing to use things they had never been shown. The guard meant to catch exactly that had a hole in it: it only recognised runs that recorded the size of the file they received, and those older runs record nothing, so they sailed through as though they were fine.
So we bought the missing measurement rather than publishing the gap or dropping the claim quietly — six fresh decks from each of the two frontier models, on the current template, scored the same way. The result is the table above: no gap. The expensive models use the catalogue as much as the cheap ones, or slightly more.
That is a smaller claim than the one we wanted, and a better one for anyone deciding whether to use this. The template is not a crutch that rescues weak models while the good ones ignore it. It is a vocabulary that every model tested actually speaks — which is the thing you would want to be true before pasting it into whatever AI you happen to have.
The layouts travel between labs
The elements were harvested from one model family’s habits. They are used by models from several others — which is the part that surprised us, because nothing in a deck tells a model that a waterfall was somebody else’s idea first.
- real table6 families
- connected graph6 families
- status chip6 families
- tiers6 families
- cycle6 families
- step exception5 families
- waterfall4 families
- coordinate plot4 families
- process chain3 families
- 2x2 grid3 families
- span bar3 families
- progress bar3 families
- native chart3 families
The sharpest single demonstration: two layouts invented one morning, from one model’s habits, were both picked up that same evening by a different lab’s model on its first exposure — for exactly the content they were designed for, in files that converted with no flattened pictures at all. The only thing that changed between the run that used them and the run that did not was whether the layout existed in the template it was handed.
What we changed
Two elements were deleted. Counting which layouts get used also shows which never do: an icon-led list and a fixed org tree had never been used by any model, in any run. Both left the template. They are kept in the test suite so either can return if the evidence changes, but a catalogue that only ever grows is a catalogue nobody has checked.
And one brief was retired from the scoring. Our own example deck answered it — the same €18 subscription split the brief describes — so a model could reach the right layout by copying rather than by choosing. That is not a measurement, and it had been quietly true for days. The example now uses different subjects, and the rule is written down: nothing in the template may answer a frozen brief.
How honest this is
Method and limits
Each rate is the share of that model’s decks that used a layout the brief called for, counted from the decks themselves rather than from anything a model reported about itself. A deck only counts if we can prove the template it received actually contained those layouts; every run that cannot prove it is excluded, which currently removes 0 decks including every frontier one. Only models with five or more provable decks appear at all.
The limits. Five or six decks per model is thin, and each brief is a single sample. “The content called for one” is our own mapping from brief to expected layout — written before the decks were scored, but written by us. A rate says nothing about whether the deck was any good; it says the model reached for a structure rather than a list of bullets. And the rates here are HIGHER than the ones we published earlier, because the excluded runs were dragging them down: this correction moved the number in our favour, which is exactly why it needed checking rather than accepting.
Nothing here involves anyone’s documents. Every deck was generated from invented briefs, and unPaper itself still sends no document anywhere.