The answer was valid JSON, which is exactly why the formula broke
Our AI returns its answer as JSON with LaTeX inside it. Nothing errored, and the formula still arrived destroyed — because \frac is valid JSON that decodes to a control character.
A student asks the AI Professor to explain a formula. The answer comes back with the formula in pieces: a stray control character, then the letters that used to be part of a command. Nothing errors. Nothing warns. The JSON parsed perfectly.
That last part is the whole problem.
The answer is not prose
When you ask a question, the app does not only want the answer. It wants three things: the answer itself, a few suggested follow-up questions, and which sources were used. So the reply is a JSON object:
{ "content": "...", "suggestedFollowUps": [ ... ], "sources": [ ... ] }
And because this is an engineering study app, content is Markdown full of LaTeX. The prompt demands it, in a section with its own heading:
=== MATH & FORMULA FORMATTING (MANDATORY) ===
Formulas are written as LaTeX, and the alternative is banned outright — a later line forbids Unicode glyphs like ∫, √, ∑, × and α, on the grounds that they render inconsistently. Every formula, in every subject, is \alpha and \frac and \nabla.
Which puts two languages inside one string, and they do not get along.
They collide on exactly five letters
JSON has a small set of single-character escapes: \", \\, \/, \b, \f, \n, \r, \t and \u.
Five of those are letters, and those same five letters are how a great many LaTeX commands begin.
So the string "\frac{a}{b}" is valid JSON. It is not a syntax error. It parses cleanly — and it decodes to a form-feed character followed by the text rac{a}{b}. The formula is gone, and every layer of the stack reports success.
Why the obvious fixes do not work
"Just parse it and catch the error." There is no error. The corruption happens inside a successful decode, so there is no exception to catch and no flag to branch on. You cannot detect a failure that is indistinguishable from success.
"Escape every backslash before parsing." That destroys the escapes doing real work. Models emit genuine \n for line breaks, and we ask them to. Quotes arrive as \". And a degree sign written as ° would stop being a degree sign and start being the literal text ° on screen.
"Tell the model to escape correctly." We do. A section headed === JSON + LATEX ESCAPING (CRITICAL) === spells out the right form and the wrong form side by side, and warns that single backslashes silently destroy the formula. The instruction is correct, and the model still does it — which is why there is a test file whose entire job is to feed the parser payloads containing single backslashes.
Prompting reduces the frequency. It does not remove the class of bug.
So we decide, one backslash at a time
The only correct move is to look at what follows each backslash and decide whether it begins a JSON escape or a LaTeX command — which means knowing LaTeX’s command vocabulary.
The repair is a single-pass scanner, and it only operates inside string literals, because that is the only place JSON escapes exist:
- A doubled backslash is copied untouched, so a correctly escaped
\\fracpasses through byte for byte. \ufollowed by four hex digits is a real Unicode escape and is preserved.- Everything else is examined. The scanner takes the longest run of letters after the backslash —
frac,nabla,times— and looks it up.
That lookup table is the interesting artefact. It is a hand-written list of 66 LaTeX commands, grouped by first letter, covering exactly the five letters JSON cares about: bar, because, begin, beta; fbox, frac; nabla, neq; rho, right; tan, theta, times.
If the run is in the list, the student meant LaTeX and the backslash is doubled. If it is not, it was a real escape — \n is a newline — and it is copied as-is. A backslash followed by something that cannot be a JSON escape at all, such as \alpha or \sqrt, is doubled unconditionally.
The repair is checked, not trusted
A scanner like this is exactly the kind of code that quietly makes things worse, so its output is never used on faith. The repaired text is handed back only if it parses. If it does not, the candidate is discarded and the next extraction strategy is tried instead.
Where the list is still wrong
Honesty requires the next part: a curated list of 66 is not LaTeX, and it has gaps we know about.
It is case-sensitive where LaTeX is not. The scanner collects letters in either case, but the list stores lower-case names. \rvert is handled. \rVert — amsmath’s double-bar norm — produces the run rVert, misses the list, and decodes to a carriage return followed by the text Vert.
It cannot be complete. Any command starting with b, f, n, r or t that nobody thought to add fails in exactly the way the routine exists to prevent. \bigoplus and \bigotimes are real operators in the digital-electronics subjects this app covers, and neither one is in the list.
And the discriminator is genuinely ambiguous. Because the letter run begins at the escape letter, a newline followed by a word is safe — \nto be precise reads as the run nto, which is not a command. But a real newline followed by u and then a non-letter is not: \nu (initial velocity) reads as the run nu, which is a command, so the line break becomes a Greek letter. LaTeX wins that tie deliberately, and it is the wrong answer some of the time.
One feature, two response contracts
One more decision that looks inconsistent until you know why.
The same feature has a streaming mode, where the answer is drawn as it arrives. That mode refuses JSON entirely and asks for plain Markdown.
The reason is this bug class. A streamed response is rendered chunk by chunk, and a half-finished JSON object cannot be parsed — a student would watch the envelope assemble itself on screen. Asking for Markdown removes the escaping collision along with the envelope.
Why not write a real LaTeX tokenizer?
Because it would not settle the ambiguity, and it would not remove the list. A tokenizer still needs a table of command names to know that \nu is a command and \nto is not, and it would still have to pick a reading when both are valid. What we would gain is a larger table and a slower parser.
The 66 names are not a compromise we are happy with. They are a compromise we can see — in one file, with a test file pointed at it — which is better than a silent corruption with no error anywhere to catch it.
Keep reading
Related reading
We asked for twenty-five questions and reported success at six
AI Quiz lets you choose 10, 15, 20, 25 or 50 questions. The parser takes that number as an argument and never reads it, so six questions for a twenty-five question request is reported as a success.
An empty list meant two different things, and we could not tell them apart
'Frequently asked' is a counting problem: read the question papers, count the repeats. The hard part is that 'nothing repeats' and 'we failed to read a single paper' both arrive as an empty list.
Every rule that makes a timetable a timetable is English text in a prompt
We went looking for the scheduler in our AI study timetable and there is not one. No overlap detection, no chronological sort, no feasibility check — the plan is whatever the model returns, and the rules exist only as sentences asking it to behave.