ununPaper
← Lab

31 July 2026

Five structures we could not name, and one improvement we measured and threw away

Synthetic narration · 6 min

When we widened our sample from one family of AI model to seven, it did not just show us what was missing from our template. It told us that two of the things we had already added were argued from too little evidence. This is what we did about that, and about a converter improvement that looked obviously right until we measured it.

Being corrected by a bigger sample

Our first round of this study ran on one model family, and two elements came out of it. A bar that spans part of a scale, because that model kept hand-building one for a plan across four quarters. And a waterfall, because asked where every euro of a price goes, it drew one three times out of three.

Both held up less well against six more models. Not one of the other families answered the four-quarters brief with a spanning bar — they all drew a grid of periods against workstreams. And where the money goes produced a different answer again: four of six built a single bar split into shares that add to a hundred percent, which is not a waterfall at all. A waterfall is a bridge, where each column starts where the last one ended. A ribbon is a split, where the total never moves.

The tempting fix is to soften the claim in the documentation. We think the honest fix is to build the thing models actually reach for and then say plainly where the line falls between the two, so the choice stops being a coin flip.

The structures this round answered, and how many model families built each
  • Period × workstream roadmap grid5 of 6
  • Layered stack with components inside each tier5 of 6
  • 100 % composition ribbon (share bar)4 of 6
  • Paired journeys on a shared scale4 of 6
  • Data donut / ring chart with a centre total2 of 6

The one we had forbidden ourselves

The last of the five is the one worth pausing on. A ring — a single number with an arc drawn around it showing its share — barely appeared in our evidence at all. That looked like a finding about taste until we traced it back to our own instructions, which had been telling every model that curved arcs would be flattened into a picture. That stopped being true when the converter improved, and nobody told the models. A structure can be missing from your data because you forbade the ingredient it is made of.

What we changed

Five elements, each added only after it converted into real, editable PowerPoint objects rather than a flat image: the roadmap grid, the parts inside each tier of a layer stack, the share ribbon, two journeys compared on one shared scale, and the ring. Every one of them was run through the conversion harness, which renders the file, converts it, opens the result and compares the two images pixel by pixel.

How closely the converted PowerPoint matches the printed page, per new element
  • Roadmap grid97.28%
  • Layer stack with parts98.08%
  • Share ribbon98.1%
  • Paired journeys98.15%
  • Ring stat98.65%

The improvement we threw away

Models hatch things. A striped bar is how they say “queued”, “at risk” or “not yet”, and those stripes are the one part of such a slide that arrives in PowerPoint as a picture instead of an editable shape. So we built the fix. Twice.

The first version used PowerPoint’s own catalogue of hatch patterns, and cost nearly five points of visual accuracy — its stripe spacing is fixed, so the texture almost disappeared. The second version drew the stripes as real geometry, and it looks, to a human eye, identical to the original. It still cost between one and two points on every stripe width we tried, because the two renderers disagree along every stripe edge.

We did not ship either. Our rule is that we never trade accuracy for editability, and a measurement is not worth keeping if it only counts when it agrees with what you were hoping to build. Hatching stays as it was. What we kept instead is the number, written into the test suite, so the next attempt has something to beat — and one useful discovery from the failed version, about where a repeating stripe pattern is anchored, which was worth five points on its own to whoever tries next.

What this means if you use unPaper

Hand your AI a brief that needs a plan across quarters, a price broken into where it goes, a before-and-after of the same two weeks, or a single figure that carries a slide, and there is now a name for the shape it wants to draw — which means it stops rebuilding one by hand, and the result arrives in PowerPoint as objects you can drag, recolour and retype.

How honest this is

Method and limits

The counts above come from 165 slides built by 7 model families against a fixed set of briefs, each brief describing content that demands a structure without ever naming one. The briefs do not change between rounds, so a model tested later stays comparable to one tested earlier. The conversion scores come from the same harness that guards every release: 77 decks, currently averaging 97.8%.

The limits. Naming the structure a model drew is a judgement made by reading the rendered slide; the counts and conversion measurements around it are not. Most briefs were sampled once or twice per model, which is enough to steer a decision and not enough to settle an argument — the two corrections above are exactly what a thin first sample costs, and the same could yet happen to something in this entry. The rendering comparison uses one office suite as its stand-in for PowerPoint, so it catches geometry and colour better than it catches type.

Nothing here involves anyone’s documents. Every slide was generated from invented briefs, and unPaper itself still sends no document anywhere.