Aabha AI Academy All articles

When RAG helps and where it can fail

Retrieval can give an AI answer relevant source material. You still need to check which evidence was found and how the answer used it.

A student asks a campus assistant, “Can thirty people use the Cedar study room?” The assistant receives this fictional current room note:

Cedar accommodates up to twenty-four people. Elm accommodates up to sixteen. For a group larger than twenty-four, contact facilities to discuss another space.

It answers, “Yes. Cedar is suitable for study groups,” and cites the note.

Identify the failure stage and cite the detail that proves your diagnosis. Write a corrected answer to the student, then explain whether retrieving more room notes would directly fix the demonstrated error.

The required limit was in the supplied note, so this example shows a failure to use the evidence. A corrected answer says Cedar's limit is twenty-four and directs a group of thirty to facilities to discuss another space. More retrieved notes alone would not address the demonstrated contradiction. The citation makes the source traceable; it does not make the claim supported.

Retrieval-augmented generation, usually shortened to RAG, combines finding external material with generating an answer from the supplied context. Lewis and colleagues' original RAG paper studied a generator combined with retrieved passages from an external index. The room example here is a simplified teaching scenario, not a result from that paper. Original RAG research

By the end, you will be able to identify whether a poor answer calls for better sources, better retrieval or better use of evidence. No code is required.

Follow the evidence through the system

In a typical application, the process has several decisions:

  1. Choose which documents the application may use.
  2. Make those documents searchable, retaining useful source and version information.
  3. Use the question to retrieve relevant passages.
  4. Give selected passages and the task instructions to a language model.
  5. Return an answer with evidence the user can inspect.

That is a common application pattern, not a requirement that every RAG system use the same search method. Search may use keywords, embeddings that represent aspects of meaning, or a combination. The important question is whether the evidence needed for this user's task reaches the answering step.

Simply retrieving a document and placing its text in a prompt does not retrain the generator's weights. It supplies information for that request. An application can refresh its searchable sources independently, but it must actually update the stored material and retrieval index before claiming current coverage.

Use RAG when the answer depends on a body of material

A campus handbook, equipment manual or changing internal reference collection can make retrieval useful. The answer may depend on details outside the model's training, and the reader may need to open the exact passage behind it.

Start by asking whether the task needs generation at all. A room-capacity filter might answer “Which rooms fit thirty people?” more directly than a conversational model. A question about one short supplied note may need careful reading without a separate retrieval system.

These are design choices, not a ranking of technologies. RAG earns its place when finding and combining relevant source material helps solve the actual question.

Diagnose three different failures

Return to the study rooms. Consider these cases:

What the assistant receives What it answers What to investigate
Only Elm's opening hours “Cedar can fit thirty people.” The needed capacity evidence was not retrieved.
An obsolete note saying Cedar fits forty “Cedar can fit thirty people.” The retrieved source is outdated for the decision.
The current note saying Cedar fits twenty-four “Cedar can fit thirty people.” The answer contradicts available evidence.

The fixes differ. Adjusting the answer prompt cannot make an obsolete source current. Adding documents cannot guarantee the model will preserve an exception already in context. First inspect the evidence that reached the model, then compare the answer with it.

RAG evaluation research similarly separates retrieval quality, faithful use of passages and generated-answer quality. Those dimensions help locate a failure; a single overall score can hide it. Ragas research

A better answer to the opening question is: “The current note limits Cedar to twenty-four people, so it does not support a group of thirty. It directs larger groups to facilities to discuss another space.” It retains the limit and the available next step without inventing a room or a booking.

Do not confuse a larger context with better evidence

More text can introduce unrelated details, contradictions and old versions. A passage may also lose its meaning if separated from its heading or an exception immediately below it.

Keep the information needed to interpret the rule together. For the room note, that includes the capacity, room name and instruction for larger groups. Retain the document's date or version when freshness matters.

Liu and colleagues found that the position of relevant information affected performance in the long-context tasks and models they studied. That result is a reason to test how your chosen model uses context, not a claim that every current model fails in exactly the same way. Lost in the Middle

Suppose a longer-context version now quotes a room's capacity correctly, but it was tested only on questions whose answers were present. Before calling it an improvement, propose one question that requires preserving an exception and one whose answer is missing. For each, state what a passing answer would do. Compare the versions on the same evidence and questions. For example, a question about a group of thirty should retain the instruction to contact facilities rather than invent a suitable room. A question about next Tuesday's availability should identify the missing schedule information rather than confirm a booking. An improvement must handle these cases as well as direct capacity questions.

Treat access and instructions as application decisions

An assistant should retrieve material the current user is allowed to access. In a real campus system, public room rules and private staff notes may need separate access controls. Apply those controls in the application rather than relying on the model to decide which material a user should see.

Retrieved text can also contain instructions aimed at changing the assistant's behaviour. NIST describes indirect prompt injection through material an application may retrieve. Source content needs to be treated as evidence to inspect, not automatic authority to perform actions. NIST Generative AI Profile, information security

A sentence in a document cannot authorize an assistant to book a room, send a message or reveal another user's record. Keep those permissions and actions explicit in the application. Prompt wording alone does not establish that the boundary works.

Try a question the note cannot answer

Using only the current fictional room note, answer: “Can I book Cedar for fifteen people next Tuesday at two?” Separate what the note establishes from what it leaves unknown. Name the evidence or permission needed to confirm a booking, then decide whether the assistant should answer, ask for evidence or perform an action.

A supported response is: “Fifteen people is within Cedar's stated capacity, but this note does not establish availability next Tuesday at two or confirm a booking.” To proceed, the assistant would need current availability for that date and time and the application's permission to make a reservation. It can explain the capacity now; it cannot confirm or perform the booking from this note alone. If your answer says yes without qualification, you have turned a capacity fact into an availability claim. If you request another capacity note, check whether it would resolve the missing schedule information.

For a small evaluation set, include questions with direct answers, missing answers, exceptions, obsolete documents and conflicting passages. Keep the retrieved evidence with each response. You can then explain what failed and whether a proposed change fixed that failure.

Continue with Before you build an API, write the request contract to define what an application accepts and rejects before connecting it to a model.