One VTU logoOne VTU
All posts
EngineeringAIRAG

Our AI could not answer a single question about your syllabus

The search was working perfectly. It just never found anything — because truncated embeddings are not unit vectors, and the threshold assumed they were.

16 July 20264 min read

The AI Professor feature is supposed to answer questions from your own course material. You index your syllabus, ask it something, and it answers from the document, telling you which module and page the answer came from.

It shipped. It answered questions about engineering in general. Asked anything about the actual syllabus sitting in its index, it would say it could not find it in your materials.

Which was, at least, honest. The retrieval was returning nothing at all.

Nothing was obviously broken

The document had been indexed. The text had been extracted, split into chunks, embedded and stored. The query was embedded the same way. A vector search ran over the chunks and compared them by cosine similarity.

Every step worked. The whole thing returned zero results.

The search had a similarity threshold — a minimum score below which a chunk is considered irrelevant and dropped. It was set at 0.3, which is a reasonable-looking number for cosine similarity between a query and a relevant passage.

Nothing was clearing it.

Truncated embeddings are not unit vectors

Here is the part that took a while to see.

The embedding model returns vectors at a default size, and supports asking for fewer dimensions. We asked for 768. That is a legitimate, documented option: the model is trained so that the first n dimensions of its output are themselves a usable embedding — smaller, faster, slightly less precise.

What that option does not do is re-normalise the result. A full-size embedding comes back as a unit vector, length 1. A truncated one does not. It is a slice of that vector, and a slice is shorter than the whole.

Cosine similarity is a dot product of two unit vectors, which is what makes it a clean number between −1 and 1. Feed it vectors that are not unit length and the arithmetic still runs — it just does not mean what you think. Similarity scores come out systematically smaller, across the board, for everything.

Real matches were scoring 0.2. Our threshold was 0.3. Nothing was relevant, because nothing could be.

The fix

Normalise every embedding to unit length the moment it comes back from the model — both when indexing a document and when embedding a query. One function, applied on both paths, so the two can never disagree about what a similarity score means.

The index had to be rebuilt afterwards, since every stored vector was the wrong length.

But one fix was not enough

Normalising fixed the scores. It did not make retrieval good, and it left a related weakness: a hard threshold is a cliff. Every chunk below the line vanishes, so a question phrased a little differently from the document's own wording can score just under it and get an empty result — even though the answer is right there.

Four changes went in alongside:

  • A lower threshold with a guaranteed fallback. The cutoff dropped to 0.15, and if nothing clears it the search still returns the best chunks it has. If your document contains anything at all, you get results rather than a shrug. Returning the least-bad passages and letting the model decide is better than returning nothing.
  • A keyword score blended in. Pure vector search is weak on rare exact tokens — a subject code, a formula's name, an abbreviation. The final ranking is 0.7 cosine similarity and 0.3 keyword overlap, so an exact term match can lift a chunk that a vector comparison alone would have missed.
  • More results. The cut went from 8 chunks to 18. Context is cheap next to a wrong answer.
  • De-duplication and better labels. Overlapping chunks are collapsed, and each source is labelled with its module and page.

Chunking for the documents we actually have

How you split a document decides what can be retrieved from it.

VTU syllabi are not prose. They are module headings, unit numbers, course outcomes and prescribed textbook lists. Splitting them into fixed-size blocks of a thousand characters cuts straight through that structure: a chunk that ends halfway through Module 3 and begins inside the textbook list is not really about anything, and it will not match a question about either.

The chunker now understands the conventions these documents use. It recognises module, unit and chapter headings, course outcomes, and the textbooks and reference books sections, and keeps each one whole. Documents without that structure fall back to overlapping sentence windows.

Every chunk carries its subject, module and page along with the text — which is what makes the citation possible when the answer comes back.

The answer still has to admit what it does not know

Better retrieval does not remove the need for the model to be honest about its sources. The prompt answers from retrieved material first and cites where it came from. When the answer genuinely is not in your documents, it says so before anything else, and only then offers a general answer explicitly marked as coming from general knowledge rather than your materials.

That distinction is the whole feature. An answer with nothing behind it is exactly the answer you should not be revising from.

What we took from it

The failure was not a crash or a wrong answer. Every component did what it was documented to do. The bug lived in an assumption shared between two of them: one side quietly returned non-unit vectors, the other quietly assumed unit vectors. Nobody was wrong, and the result was silently empty — which is much harder to notice than an error.

This, in the app

The part of One VTU this post is about.

Keep reading

Related reading

EngineeringAIProduct

Every rule that makes a timetable a timetable is English text in a prompt

We went looking for the scheduler in our AI study timetable and there is not one. No overlap detection, no chronological sort, no feasibility check — the plan is whatever the model returns, and the rules exist only as sentences asking it to behave.

12 September 20266 min read