An empty list meant two different things, and we could not tell them apart
'Frequently asked' is a counting problem: read the question papers, count the repeats. The hard part is that 'nothing repeats' and 'we failed to read a single paper' both arrive as an empty list.
Top Questions is a counting problem disguised as an AI feature. Find the previous-year question papers for a subject, read them, and count which questions come up more than once.
The counting is easy. The hard part is that “the questions never repeat” and “we failed to read a single paper” are the same empty list.
Almost everything interesting in this feature is machinery for telling those two apart.
Where the papers come from
Papers live in one table, filtered by scheme, branch and semester, and matched to the subject on either of two columns — the semester-one subject code or the semester-two one — because a subject can sit in either slot.
They come back newest first, keyed by a session label like dec_jan_2024. Each paper is downloaded, turned into text, and prefixed with its session so the model can cite years rather than just count:
--- PYQ: Dec-Jan 2024 ---
Newest first matters more than it looks, and we will come back to it.
Reading a PDF that might be a photograph
These are university question papers. Some are real PDFs with a text layer. Some are scans of a photocopy of a printout.
So extraction tries the cheap thing first: pull the embedded text. If it yields at least 50 characters, that is the paper. Anything less, and the code assumes it is a scan — it renders every page to a bitmap and runs OCR over it.
The OCR runs on the device, through the platform’s on-device text recognition. None of it is uploaded, which matters here more than anywhere else: this stage happens before a single byte goes to a model, so the papers are read on the phone and only the extracted text ever leaves it.
The ambiguity, and what we do about it
Follow the failure paths and they all converge on the same value.
- A download that fails is skipped. If every download fails, the concatenated text is empty.
- A paper whose text extracts to nothing is skipped, for the same result.
- A profile whose USN is too short to read a branch code from returns empty before any fetch happens.
The prompt is then handed an empty block and correctly returns no questions. Which is indistinguishable from a subject where nothing repeats.
So the result type carries an explicit flag for “analysed, and the answer really is nothing”, separate from a failure. The screen renders them as different sentences, and there is a third state for “not analysed yet” shown when there is no cached result at all.
That is why the “no questions” message is specific rather than generic. A student reading “no question was asked twice or more across the available papers” is being told something true. A student reading “something went wrong” would not know whether to trust it.
We chose to cache the empty answer, and here is the price
An empty result is a result. It is written to the cache like any other, and the reason is cost: the pipeline downloads and OCRs every paper for every subject, and re-running it on every visit would be absurd.
The price is exact and worth stating plainly. A bad extraction that produces an empty answer is cached as “nothing repeats” and stays that way, because the cache has no expiry and no version. The only ways out are a manual re-sync or a failure, which clears that subject’s entry.
That last part is deliberate. On any failure the subject’s cached entry is deleted, so a stale green tick cannot survive a re-sync that did not actually happen. A cached success and a cached failure should not be able to coexist.
What we do not enforce
Three things, and they are all visible in the app if you look.
The “asked twice or more” rule is prompt text only. The prompt says to include only questions with a count of two or more. The parser does not check it. A question the model returns with a count of one is parsed, cached, and displayed as “Asked 1x” — inside a feature whose entire premise is repetition.
A missing count becomes zero. The field is parsed with a fallback to zero, so a malformed response shows “Asked 0x” rather than being rejected.
Module names are never normalised. Grouping is done on the raw string the model returned, which means Module 1, 1 and Module 1: Calculus become three separate sections for the same module, in whatever order the model happened to emit them. Only the questions inside a section are sorted, by count.
What we pay for and then throw away
Every paper is downloaded and, if it is a scan, rendered and OCR’d page by page. Each one is prefixed with its session label. Then the concatenation is cut at 30,000 characters with a plain substring.
Nothing is wasted in the common case, because papers arrive newest first and the newest are the most useful.
But for a subject with many sessions, the older papers were still fetched and still OCR’d, and their text never reaches the model. We did the expensive part for output nobody sees. A per-paper budget would fix it; we have not written one.
And nothing tells the model — or the student — when a paper was dropped. Five sessions in, two unreadable, and the prompt still reads === PREVIOUS YEAR QUESTIONS === with no indication that a third of the evidence is missing.
The guard rails this feature does not have
This one is the odd feature out, and it is worth comparing to its neighbour.
A full sync runs every subject in the profile, one after another, with no cancel button. Each provider attempt gets a 60-second timeout, and a failing provider triggers a sweep across every candidate model and every configured key. There is no overall deadline of any kind.
The AI Study Timetable — built by the same team, on the same failover engine — caps a whole generation at three minutes, with a comment explaining that otherwise a provider which is slow rather than dead can keep a student waiting for many minutes. Top Questions kept the engine and dropped the ceiling.
It also does not go through the app-wide queue that keeps background AI work from stampeding a provider, and it does not cost a credit. Which means the most expensive operation in the app is the one with the fewest brakes on it.
Why we are writing this down
Because it is the honest state of the feature, and the parts that are missing are the parts a student would want to know about.
What works is the part we spent the effort on: reading papers that are sometimes photographs, entirely on the device, and being careful never to present “we could not tell” as “there is nothing there”. That distinction is the whole feature.
What is missing is enforcement of the rule the feature is named after, and a ceiling on how long a Sync can run. Both are on the list.
Keep reading
Related reading
A topic you have not opened in a month still says 'Yesterday'
Daily Revision schedules reviews on a five-rung ladder and then labels them with times that cannot be right. An item untouched for a month still reads "Yesterday".
We asked for twenty-five questions and reported success at six
AI Quiz lets you choose 10, 15, 20, 25 or 50 questions. The parser takes that number as an argument and never reads it, so six questions for a twenty-five question request is reported as a success.
Every rule that makes a timetable a timetable is English text in a prompt
We went looking for the scheduler in our AI study timetable and there is not one. No overlap detection, no chronological sort, no feasibility check — the plan is whatever the model returns, and the rules exist only as sentences asking it to behave.