The image model cannot spell, so Flutter draws the labels
We asked a diffusion model for a labelled diagram and got one back with PISITON written on it. Every label a student reads in this feature is now painted by Flutter, and the image is told not to draw text at all.
Ask an image model for a diagram of an electric motor and it will produce something that looks, at a glance, exactly right. Then you read the labels.
Somewhere on the housing it says PISITON.
That is a real example, from our own code comments, written down the day we stopped asking diffusion models to draw text.
Why we were asking at all
Learn with Gen AI takes one typed topic and builds a lesson from it: an explanation, a diagram, and then a quiz or flashcards if you want them. The diagram is the part people screenshot, so we wanted it to be a real illustration rather than a box of bullet points.
The obvious pipeline is: ask a model for the lesson and a picture prompt, then hand that prompt to an image model. Which works beautifully until the picture has words in it.
The rule we ended up with
Diffusion models are extremely good at the shape of text and extremely bad at the spelling. They render something that has the visual rhythm of a word, in the right place, in a plausible font — and it is not the word.
For a product whose entire value is teaching engineering accurately, that is not a cosmetic problem. A diagram labelled PISTION is worse than no diagram, because the student believes it. A wrong answer they will question; a wrong label on a confident-looking illustration they will not.
So the rule is absolute: the image model is never allowed to draw text.
It is enforced in one place so that no call site can forget it. Every image prompt ends with the same appended paragraph:
ABSOLUTE REQUIREMENT: Do not render labels, captions, text,
numbers, arrows, letters, digits, watermarks, legends, titles
or lettering of any kind anywhere in the image.
Sentences are weak; negative prompts are not
That paragraph is not enough on its own. Diffusion models are trained to follow instructions loosely, and "do not draw text" competes with everything else in the prompt describing what should be there.
So the same rule is sent a second time, in the form these models actually obey — a negative prompt:
text, letters, words, labels, captions, numbers, digits,
typography, handwriting, watermark, signature, logo, legend,
title, annotations, callouts, arrows with text
We send it both ways because they are not redundant. A sentence asking for a clean diagram is a request; a negative prompt is a constraint. The second one is what actually keeps the lettering out.
Where the labels come from
From Flutter, at render time.
The image model returns an illustration. Separately, the reasoning model returns a list of components — the parts worth naming and pointing at. Each one arrives with an anchor position expressed as a fraction in 0–1 space, because pixels would not survive the image being scaled to the phone in your hand:
{ "label": "Stator", "anchor_x": 0.32, "anchor_y": 0.58 }
Flutter paints the label and draws the leader line to exactly that point. Every word the student reads on a diagram is drawn by the same text renderer as the rest of the app, which means it is spelled correctly, it is selectable, it scales with the system font size, it re-themes in dark mode, and it is readable by a screen reader.
None of that is true of text baked into a JPEG.
There is a cap of ten painted components per diagram, because past that the leader lines and labels overlap into an unreadable wall on a phone screen. Components beyond the cap are not lost — they appear in a "Key Components" list underneath.
Half the visuals were never pictures
The other half of the cost story is noticing that some of the things we were paying an image model to draw are not drawings at all.
A mind map is a tree of text with lines between it. An infographic is boxes and headings. Both are pure text, and Flutter draws them from structured data for nothing, correctly, and in the student’s language.
So the visual types are split in two, and the split is the spend decision:
- Drawn by Flutter — mind maps, infographics, formula sheets. These cost no image call and no network.
- Generated — technical diagrams, real-world scenes, anything that genuinely needs to be an illustration.
Each visual type carries a flag saying which side it is on, and the flag is called needsGeneratedImage. That single boolean is the money gate.
The call that decides whether to spend anything
Even with that split, there is a temptation to have the model decide by writing whatever it feels like. Ask a strong model for a lesson and it will enthusiastically produce a diagram for "explain Ohm’s law" — and we would pay for a picture nobody asked for.
So the pipeline starts with a separate, cheap classification call. It runs first, at a temperature of 0.1, and its only job is to answer questions about the topic: what kind of material is this, how deep should it go, and does a picture actually earn its cost?
Its answer is reduced to a single flag. Everything downstream reads that flag rather than the model’s enthusiasm:
- The picture spec is only parsed when the verdict said a picture was wanted.
- The spec is then checked against the plan, so a spec that disagrees cannot quietly upgrade the decision.
- The provider only calls the image model under one condition, and that condition is the flag.
The upshot: a model that writes a diagram nobody planned is ignored outright. Its spec is thrown away. Only the thing that decided the picture was worth paying for can authorise paying for it.
If the classification call fails entirely, it is not fatal. A conservative on-device heuristic stands in — narrow on purpose, because a keyword-only guess about what "looks visual" would spend an image call on topics that merely sound visual.
No model id is hardcoded anywhere in the feature
The image models are rows in a database table, not constants in the code. The resolver looks up rows marked as image models for the configured providers, and only falls back to asking a provider for its model list when the catalogue contributed nothing.
That was not a preference, it was forced. The provider whose free tier we lean on publishes a model list that is chat-only — 130 models, none of them image models — so there was no id we could have hardcoded even if we wanted to. A row is the only way an image model becomes callable, and setting it inactive retires it without shipping an app update.
Routing names tags, never a model id: a cosmetic style and a structural type resolve to a set of capability tags, and candidates matching more of them are tried first. The rule is that this only ever reorders — no candidate is dropped, so a tag-matched model that fails still steps aside for the next one.
What we gave up
Diagrams are Hugging Face only. Chat models may come from any configured provider; the picture generator is restricted to one. That is a real limitation and it is why the failure message is specific rather than generic.
We cannot show you a labelled illustration. Everything on a diagram is Flutter text over a picture, which means the label and the thing it names are only aligned by the coordinates the model gave us. When those are slightly off, a leader line points at the wrong part of the housing — a smaller failure than PISITON, but a visible one.
We took it. A picture that cannot spell is a picture you cannot learn from, and an illustration with a Flutter label over it is one you can.
Keep reading
Related reading
Five AI workers walked into a rate limit at the same time
Our app opened, five background jobs woke up at once, and every one of them reached for the same API key. Here is the queue that fixed it — and why the chat feature is deliberately not allowed to use it.
The card you mark as known never comes back
Flashcards has three states and no spaced-repetition scheduler. Marking a card known removes it from revision permanently, and four different screens disagree about which cards are due.
The Daily Quiz has no daily content, and that is deliberate
Daily Quiz seeds nothing by date and generates nothing when the app opens. The quizzes are the output of a durable background queue, and the only daily thing about the feature is which five of them we remind you about.