ununPaper
← Lab

31 July 2026

Seven AI models, one slide brief, and the diagram none of them had a name for

Synthetic narration · 7 min

We spent a day asking seven families of AI model to build the same slides, then measured what came back — in a real browser, and again after converting each file to PowerPoint. The point was not to rank them. It was to find out what our own template was missing.

Why we ran it

unPaper hands your AI a set of rules: a fixed page, no scripts, nothing fetched from the network, and only the kinds of elements that survive becoming editable PowerPoint. Those rules are why a deck prints exactly and converts cleanly. But they carry a risk nobody ever reports as a bug: if a model would have built something richer, and it builds something plainer because our template has no name for the richer thing, the deck still looks fine. It is just quietly less than the model could do.

So we wrote a set of briefs that describe content demanding a structure without ever naming one — no “draw a funnel”, just the situation — and ran each one twice. Once with the model working alone, once through unPaper. The difference between those two answers is the part that tells you something.

What we found

The rules do not cost a strong model its structure. In eight of twelve cases, a frontier model on our contract built exactly the same structure it had built unaided. It simply had to construct it by hand, because we had not given that structure a name. That reframed the whole exercise: the signal that an element is missing is not an ugly deck, it is a model reinventing the same thing over and over.

Some structures every model reaches for. Asked how four systems depend on one another, every family we tested drew the same thing: boxes placed freely, joined by lines, with one connection drawn heavier than the rest because it carries the risk. We had a name for a chain and a name for a tree. We had no name for a graph — and that brief produced more broken, overflowing slides than any other, which is what a missing idiom looks like from the outside.

Structures built by multiple model families, which our template could not name
  • Node-edge system map6 of 6 models
  • Closed loop / flywheel6 of 6 models
  • Exception / stall annotation pinned to a step6 of 6 models
  • Layered stack with components inside each tier5 of 6 models
  • 100 % composition ribbon (share bar)4 of 6 models
  • Paired journeys on a shared scale4 of 6 models

What changed in the template

Seven new elements, each one added only after it survived the test that matters: converting into real, editable PowerPoint objects rather than a flat picture. A span bar for anything that starts partway along a scale. A waterfall. A stack of tiers. A plot on two real axes, which is not the same thing as a two-by-two grid. A graph whose arrows are genuine PowerPoint connectors — drag a box in the meeting and the arrow follows it, which is something no model produces on its own, because a static line is all a web page can express. A loop that returns to where it began. And a small flag that pins an exception to one step and carries a number, because every model marked where a process hurts and how much, and none of them had a tidy way to do it.

We also deleted two. A measurement of which elements models actually reach for showed two of ours had never been used by any model, in any run: an icon-led list and a fixed org tree. Both left the template. They are kept and still tested, so either can come back if the evidence changes, but a catalogue that only grows is a catalogue nobody has checked.

The two zeroes turned out not to mean the same thing, which is worth saying because we got it wrong first. The icon list had a worked example in the template and was still ignored — that is a real preference. The org tree had no example at all, only a name. Its zero says much less about what models want and much more about something we now believe generally: a name without a worked example does not get used.

The part we did not expect

Cheaper, smaller models use the template’s vocabulary about twice as often as the frontier ones do.

How often a model used a named element when the content called for one
  • glm-5.283%
  • gpt-oss-120b76%
  • mistral-small-3.2-24b-instruct72%
  • deepseek-v4-flash67%
  • gpt-5.250%
  • claude-opus-525%

That is the design working rather than a defect. The instructions invite a strong model to compose beyond the catalogue and a weaker one to reach into it. Which means the template does most of its work exactly where it is most needed — and it is the clearest evidence we have that the gap between an expensive model and a cheap one narrows when both are given the same structure to build in.

What we got wrong

Two of our own mistakes are worth recording. Our instructions had been telling every model that curved lines and arcs would be flattened into images — true once, and quietly false after the converter improved. Models were obediently drawing straight lines instead of curves for no reason. And a run that measured whether models used our newest elements returned a confident no, which turned out to be nonsense: the file those models received had been generated before the elements existed. Both are fixed. We publish them because a page that only prints wins is an advertisement.

How honest this is

Method and limits

154 slides across 7 model families, each rendered in a real browser with scripts disabled — the same view our converter has — and then run through the converter itself, where the result is counted as native objects, partly flattened, or broken. 84% of the corpus converted fully natively. The briefs are held fixed so that every model tested later remains comparable to the ones tested already.

The limits matter. One model family is over-represented, because the study started with it. Most briefs were sampled once or twice per model, which is enough to steer a decision and not enough to settle an argument. The names we give the structures models drew are a judgement made by reading the rendered slides; the counts, conversions and measurements around them are not.

Nothing in this study involves anyone’s documents. Every slide was generated from invented briefs, and unPaper itself still sends no document anywhere.