Skip to main content

Command Palette

Search for a command to run...

RAG(Retrieval-Augmented Generation): The Good, the Bad, and the Hallucinations

Updated
•14 min read•View as Markdown

Understanding RAG: How It Improves LLM Answers—and Where It Can Still Fail

Large Language Models (LLMs) such as ChatGPT can answer questions, write code, summarize documents, and generate content in seconds. But they have an important limitation: they do not automatically know your latest company documents, private files, product details, or newly published information.

An LLM only knows what it learned during training and what is included in the conversation. If you ask it about information outside that scope, it may guess, give outdated information, or confidently provide an incorrect answer.

This is where Retrieval-Augmented Generation (RAG) becomes useful.

RAG is a technique that gives an LLM relevant external information before it generates a response. Instead of relying only on its training data, the model can retrieve useful documents, policies, notes, FAQs, or database content and use them as context.

However, RAG is not a magic solution. It can improve answer quality significantly, but it cannot guarantee that every response will be correct.

This article explains what RAG is, how it works, where it is useful, and why it can still fail.


The Limitation of an LLM Without External Knowledge

Imagine you build a chatbot for an e-commerce company.

A user asks:

What is the return policy for electronics purchased during a sale?

A normal LLM may try to answer based on general knowledge. It may say that electronics can be returned within 30 days, even if your company policy allows only 7 days for sale items.

The model is not intentionally trying to mislead the user. It simply does not have access to the company's actual policy document.

The same issue happens when users ask about:

  • Internal company documentation

  • Product manuals

  • Customer support policies

  • Legal guidelines

  • Personal notes

  • Private databases

  • Recently updated information

  • Research papers published after the model's training period

Without external context, an LLM often has two choices: say it does not know, or generate an answer that sounds plausible.

RAG helps reduce this problem by giving the model relevant information before it answers.


What Is RAG?

RAG stands for Retrieval-Augmented Generation.

It combines two steps:

  1. Retrieval: Find relevant information from an external knowledge source.

  2. Generation: Give that information to an LLM so it can generate a better answer.

Instead of asking the model to answer from memory alone, a RAG system first searches a knowledge base.

For example, if a user asks:

How do I reset my account password?

The system can search company help documents, find the password reset instructions, and provide those instructions to the LLM. The LLM then creates a clear response based on the retrieved content.

A simple RAG flow looks like this:

User Query → Retrieval → LLM → Response

The LLM is still responsible for generating the final answer, but it now has useful context to work with.


How a Basic RAG Pipeline Works

A basic RAG pipeline usually has four main steps.

1. Store Knowledge in a Searchable Form

First, documents are collected from sources such as:

  • PDFs

  • Documentation pages

  • FAQs

  • Product manuals

  • Support tickets

  • Database records

  • Internal company wikis

  • Blog posts

These documents are usually broken into smaller pieces called chunks.

For example, a 50-page employee handbook might be divided into many smaller sections such as:

Chunk 1: Leave policy
Chunk 2: Remote work policy
Chunk 3: Expense reimbursement policy
Chunk 4: Employee benefits

Each chunk is then stored in a searchable knowledge base.


2. Convert the User Query into a Searchable Format

When a user asks a question, the RAG system tries to understand its meaning.

For example:

Can I work from home on Fridays?

The system searches for chunks related to remote work, work-from-home rules, flexible schedules, or company attendance policies.

The goal is not only to match exact words. A good retrieval system should also understand related meaning.

For example, the phrases below may all be related:

Work from home
Remote work
Hybrid work
Flexible location policy

3. Retrieve Relevant Chunks

The system selects the most relevant chunks from the knowledge base.

For the question:

Can I work from home on Fridays?

It may retrieve:

Remote Work Policy:
Employees may work remotely up to two days per week with manager approval.

This retrieved text becomes context for the LLM.


4. Generate the Final Response

The LLM receives both the user question and the retrieved context.

User Question:
Can I work from home on Fridays?

Retrieved Context:
Employees may work remotely up to two days per week with manager approval.

The LLM can then generate a response such as:

Yes, you may be able to work from home on Fridays if it fits within the two remote days allowed per week and your manager approves it.

This answer is more reliable than asking the LLM to guess the company policy.


RAG Architecture Diagram

                ┌──────────────────┐
                │   User Question  │
                └────────┬─────────┘
                         │
                         ▼
                ┌──────────────────┐
                │ Retrieval System │
                │ Search Knowledge │
                │      Base        │
                └────────┬─────────┘
                         │
                         ▼
                ┌──────────────────┐
                │ Relevant Context │
                │ Documents/Chunks │
                └────────┬─────────┘
                         │
                         ▼
                ┌──────────────────┐
                │       LLM        │
                │ Generates Answer │
                └────────┬─────────┘
                         │
                         ▼
                ┌──────────────────┐
                │ Final Response   │
                └──────────────────┘

Where RAG Works Well

RAG is especially useful when an application needs answers based on specific, changing, or private information.

Common use cases include:

Customer Support Chatbots

A support chatbot can retrieve answers from help articles, return policies, troubleshooting guides, and product documentation.

For example:

How do I cancel my subscription?

Instead of guessing, the chatbot can retrieve the exact cancellation instructions.


Internal Company Knowledge Assistants

Employees often need information from internal documents.

Examples include:

  • How to request leave

  • How to submit expenses

  • Security policies

  • Engineering documentation

  • HR policies

  • Onboarding guides

A RAG system can help employees search internal knowledge using natural language.


Software products often have large documentation websites.

A user may ask:

How do I configure OAuth in this product?

A RAG system can retrieve the relevant documentation section and help generate a direct explanation.


Professionals often work with large sets of documents.

For example:

  • Legal contracts

  • Compliance policies

  • Medical research papers

  • Financial reports

RAG can help users locate relevant sections quickly. However, these domains require extra care because incorrect answers can have serious consequences.


Personal Knowledge Assistants

RAG can also be used with personal notes, bookmarks, PDFs, or saved articles.

For example:

What did I write about my project idea last month?

The system can search your notes and provide a summary based on the retrieved content.


Good Retrieval vs Poor Retrieval

The quality of a RAG system depends heavily on retrieval.

Consider this question:

What is the refund policy for annual subscriptions?

Good Retrieval

The system retrieves:

Annual subscriptions can be refunded within 14 days of purchase if less than 20% of the service has been used.

The LLM can now produce a useful answer.

Poor Retrieval

The system retrieves:

Monthly subscriptions can be canceled at any time.

This information is related to subscriptions, but it does not answer the user's question about annual refunds.

The LLM may still generate a confident response, but the answer may be wrong because the context was wrong.

A simple comparison looks like this:

Good Retrieval:
Question → Correct policy document → Accurate response

Poor Retrieval:
Question → Unrelated document → Incorrect or confusing response

RAG quality is often limited by the quality of the retrieved information.


Why RAG Sometimes Gives Incorrect Answers

Many people assume that adding RAG automatically makes an LLM accurate. In reality, RAG improves the chances of a correct answer, but several things can still go wrong.


1. Poor Retrieval and Missing Context

The system may fail to find the correct document.

For example, a user asks:

What happens if I miss a payment deadline?

But the retrieval system finds a document about account cancellation instead of late payment fees.

The LLM may use the wrong document and create an inaccurate answer.

Sometimes the correct document exists, but the search system does not retrieve it because:

  • The query uses different wording

  • Important keywords are missing

  • The document is poorly indexed

  • The document is stored in the wrong category

  • The retrieval system returns only a few results

  • The relevant content is ranked too low

If the LLM does not receive the right context, it cannot reliably produce the right answer.


2. Poor Chunking Can Hurt Response Quality

Before documents are stored in a RAG system, they are usually split into chunks.

Chunking is important because an LLM cannot receive an entire library of documents for every question.

However, bad chunking can remove important context.

Imagine this policy:

Employees may work remotely up to two days per week.
Remote work requires manager approval.
Remote work is not allowed during the first 30 days of employment.

If this is split badly, the system may retrieve only:

Employees may work remotely up to two days per week.

The LLM may answer:

Yes, you can work remotely two days per week.

But it may miss the important conditions about manager approval and the first 30 days of employment.

A better chunk keeps related information together:

Employees may work remotely up to two days per week with manager approval.
Remote work is not allowed during the first 30 days of employment.

Chunking Comparison

Poor Chunking:
[Employees may work remotely up to two days per week.]

[Remote work requires manager approval.]

[Remote work is not allowed during the first 30 days.]

Result: Important rules may be separated.

Better Chunking:
[Employees may work remotely up to two days per week with manager approval.
Remote work is not allowed during the first 30 days of employment.]

Result: The LLM receives complete context.

Chunk size, overlap between chunks, headings, and document structure can all affect response quality.


3. Context Window Limitations

LLMs have a limited amount of text they can process at one time. This limit is called the context window.

A RAG system may retrieve many relevant documents, but not all of them can always fit into the model's context.

For example, imagine a user asks:

Compare all security requirements across our 20 internal policy documents.

The system may find useful information in many documents. But if only a few chunks fit into the context window, some important details may be left out.

Knowledge Base:
Document A
Document B
Document C
Document D
Document E
Document F

LLM Context Window:
Document A
Document B
Document C

Missing:
Document D
Document E
Document F

If the missing documents contain exceptions or updated rules, the final answer may be incomplete.

Even when a model has a large context window, sending too much information can create other problems:

  • More cost

  • Slower responses

  • Important details getting buried in too much text

  • Lower-quality reasoning due to irrelevant context

RAG is not just about retrieving more documents. It is about retrieving the most useful documents.


4. Hallucinations Can Still Happen

RAG reduces hallucinations, but it does not eliminate them.

A hallucination happens when an LLM generates information that is not supported by the available context.

For example, the retrieved document says:

Users can request a refund within 14 days.

But the model answers:

Users can request a full refund within 30 days, and processing takes 3 business days.

The model may have added details that were never present in the retrieved content.

This can happen because LLMs are trained to generate fluent, helpful-sounding language. If the context is incomplete, the model may fill gaps with assumptions.

To reduce hallucinations, RAG systems can be designed to:

  • Ask the model to answer only from the provided context

  • Show source citations with the answer

  • Tell the model to say “I don’t know” when information is missing

  • Use confidence checks

  • Verify answers against retrieved documents

  • Limit unsupported claims

Still, no prompt or system design can guarantee zero hallucinations.


5. Knowledge Bases Can Become Outdated

A RAG system is only as current as its knowledge base.

Imagine a company updates its refund policy from 30 days to 14 days. If the old document remains in the knowledge base, the system may retrieve outdated information.

This can lead to confusing or incorrect answers.

Knowledge bases need regular maintenance:

  • Add new documents

  • Remove outdated documents

  • Update changed policies

  • Track document versions

  • Re-index content after updates

  • Mark old content as archived

For fast-changing information, such as product pricing, stock availability, regulations, or live system status, a static RAG knowledge base may not be enough.

In these cases, the system may need direct access to live APIs or databases.


When RAG Is Not the Right Solution

RAG is useful, but it is not the best solution for every problem.

When You Need Real-Time Data

If users ask questions like:

What is my current account balance?

or:

Is this product in stock right now?

A RAG system may not be enough because these answers change frequently.

A direct database query or API call is usually better.


When You Need Calculations

If the task requires precise calculations, RAG alone is not ideal.

For example:

Calculate the total tax for these 500 transactions.

The system should use a calculator, code execution tool, or financial system instead of relying only on retrieved text.


When You Need Strict Accuracy

Some tasks require deterministic and verifiable answers.

Examples include:

  • Bank transfers

  • Medical diagnosis

  • Legal decisions

  • Tax filing

  • Security permissions

  • Database updates

RAG can help retrieve relevant information, but the final decision should not rely only on an LLM-generated response.


When the Knowledge Base Is Small and Structured

If you have a small database with clearly structured fields, traditional search or database queries may be simpler.

For example:

Find all orders placed by user ID 123.

A database query is faster, more accurate, and easier to validate than using RAG.


RAG Improves Answers, But Does Not Guarantee Correctness

RAG is powerful because it gives LLMs access to information beyond their training data.

It can make chatbots more useful, help employees search internal documentation, improve customer support, and provide answers based on private or recent knowledge.

But RAG systems can still fail when:

  • The wrong documents are retrieved

  • Important context is missing

  • Documents are chunked poorly

  • The model reaches context window limits

  • The model hallucinates details

  • The knowledge base is outdated

  • The task requires live data or exact calculations

The key idea is simple:

RAG improves the information available to an LLM, but it does not guarantee that the LLM will always interpret or use that information correctly.

A strong RAG system needs more than an LLM and a document search feature. It needs good data, good chunking, accurate retrieval, updated knowledge sources, and careful evaluation.

When used for the right problems, RAG can make AI applications much more useful. When used without understanding its limitations, it can still produce answers that sound confident but are incomplete or wrong.