We asked for twenty-five questions and reported success at six
AI Quiz lets you choose 10, 15, 20, 25 or 50 questions. The parser takes that number as an argument and never reads it, so six questions for a twenty-five question request is reported as a success.
AI Quiz lets you pick how many questions you want: 10, 15, 20, 25 or 50. Pick 25 and you might get six. The app will tell you the quiz is ready, score you out of six, and show you a percentage. Nothing on the screen looks broken, because every number the screen shows is derived from whatever came back.
The parameter we passed and never read
The parser takes the count you asked for:
List<QuizQuestionModel> _parseQuestions(String raw, int expectedCount)
expectedCount appears exactly twice in that file: in the signature above, and at the single call site that passes it. It is never mentioned again inside the function body. The body does its own work — accept a bare array or the first of quiz, questions, data or results, map each entry into a model, drop any whose question text is empty, re-index the ids — and then returns.
Success is decided elsewhere, and it is generous:
bool get isSuccess => questions.isNotEmpty && errorMessage == null;
One question is a success. Six questions for a twenty-five question request is a success. The count was threaded through the whole path for the purpose of checking it, and then not checked. Every denominator in the player and the result screen comes from the questions that arrived, so the mismatch is invisible from the inside.
We had braced for a different failure
The prompt asks for LaTeX inside JSON strings and requires every backslash to be double-escaped:
"Every LaTeX backslash MUST be double-escaped (\\\\) inside JSON string values"
That instruction exists because models were not doing it. \frac, \theta and \% are not legal JSON escapes, so jsonDecode throws, and a complete quiz is lost to one character. So we wrote a repair. Strict parsing is tried first, and only on a FormatException do we fall back to a pass that doubles every backslash which is not part of \\, \", \/ or \uXXXX.
The strict-first order is deliberate, and the comment in the code says why: \t, \n and \f are valid JSON escapes, so escaping them blindly would corrupt real newlines and tabs in valid input.
That reasoning has a price we accepted without naming it. \theta begins with \t. \frac begins with \f. \beta begins with \b. Strict parsing therefore succeeds on those payloads, the repair never runs, and the command is quietly eaten into a control character. The answer parses, the question arrives, and the formula is subtly wrong. We are describing the mechanism the comment names, not a measured incident — we have not counted how often it happens, and the code carries no log that would tell us.
The failure class with no defence at all
Everything above is about text that will not parse. There is a second class — JSON that parses perfectly and is not a quiz — and almost nothing in the path looks at it.
A question with no options. A question with three. An answer value that is not the text of any of its options. The prompt demands "exactly 4 options" and then, a few lines later, its own worked example shows three:
"options": ["$r^2$", "$r$", "$1$"]
There is a validator on the model. It reads:
bool get isValid => options.length == 4 && answer.isNotEmpty;
It is never called. Not by the parser, not on the way into the player, not before scoring. The only filter that runs is the empty-question-text check.
The comment that promises something the code does not do
The model's own doc comment says options "may arrive as List or as a JSON-encoded string". The line beneath it is:
json['options'] is List ? json['options'] as List : <dynamic>[]
A string-encoded options array becomes an empty list. The question survives, counts toward the total, and renders zero tappable tiles. That student's ceiling on that quiz is permanently below 100%, and nothing on any screen says why.
Scoring is exact string equality against the answer field, in three separate places. If a model returns "B" where the options are real text, no answer on that question can ever be marked correct — the review screen just prints the correct answer with no option ticked beside it. The Learn screen is the one scorer that defensively trims before comparing, and it carries a note saying why.
What it costs when it fails
A quiz credit is spent immediately before the call, and there is no refund path anywhere in the credit system. The comment beside it is honest about the timing and wrong about the protection: spending at the point of generation does mean a credit is never spent on a quiz you never asked for, but a malformed reply still costs one.
Failover also treats a malformed reply differently from a rate limit. Only a rate limit walks the provider's remaining keys; anything else charges the whole model slot and moves to the next provider. So one bad generation costs more than one key.
Why not just ask the model to fix it?
Because we already paid for the text. A "your JSON was invalid, repair it" round trip doubles the latency and the quota to recover something we can fix locally in a few lines, and the repair we have handles the common case without a second call.
The genuine gap is truncation. Ask for 50 questions, get 38 and an unclosed bracket, and the extractor returns null and the parser returns an empty list. An 80% complete quiz is thrown away whole. Recovering the complete objects from a truncated array is the change we would make first, followed by actually reading expectedCount and calling the validator that already exists. Neither is done.
Keep reading
Related reading
The answer was valid JSON, which is exactly why the formula broke
Our AI returns its answer as JSON with LaTeX inside it. Nothing errored, and the formula still arrived destroyed — because \frac is valid JSON that decodes to a control character.
An empty list meant two different things, and we could not tell them apart
'Frequently asked' is a counting problem: read the question papers, count the repeats. The hard part is that 'nothing repeats' and 'we failed to read a single paper' both arrive as an empty list.
Every rule that makes a timetable a timetable is English text in a prompt
We went looking for the scheduler in our AI study timetable and there is not one. No overlap detection, no chronological sort, no feasibility check — the plan is whatever the model returns, and the rules exist only as sentences asking it to behave.