LLM 4 - Prompt Engineering and RAG
Welcome back. In LLM 3 we learned how decoding, demonstrations, reasoning patterns, tools, and post-processing shape an LLM application. Today we will make one distinction that changes how we design the entire system:
A better prompt can clarify a task, but it cannot manufacture missing evidence.
Suppose a customer asks whether product BB43300 has a two-year warranty. If no applicable policy has been supplied, rewriting the question ten times will not reveal the answer. The application needs both a task specification and evidence. Prompt engineering controls how the model should use information; retrieval supplies the information that may support the answer; validation checks whether the answer actually follows from it.
This lesson develops that full path. We begin by turning prompts into testable specifications, then build a small support assistant whose answers must cite a policy or abstain. From there we construct a retrieval-augmented generation system one component at a time: inverted indexes and BM25, dense embeddings, chunking, approximate nearest-neighbor search, NSW and HNSW graphs, hybrid retrieval, reranking, context construction, citations, and failure diagnosis. By the end, “RAG” will no longer mean “attach a vector database.” It will mean a sequence of separately testable decisions about what evidence reaches the model and what the model is allowed to claim.
Prompts Become Testable Specifications
A Prompt Is More Than a Question
A prompt is any input that asks a model to produce a desired output. It may be a short instruction—“Describe a holiday destination”—or a small program written in natural language, with rules, data, examples, and an output schema. The second view is more useful for serious applications because it forces us to ask what information belongs to the task and what success means.
Four building blocks provide a practical starting point:
| Block | Question it answers | Example in a support system |
|---|---|---|
| Instruction | What should the model do? | Answer a warranty question using policy evidence. |
| Context | What setting and rules matter? | Product, region, effective date, and citation policy. |
| Input data | What material should be processed? | The user question and selected policy passages. |
| Output requirement | What form and behavior count as success? | Return an answer, evidence IDs, or insufficient evidence. |
The blocks are related but not interchangeable. An example answer is not current evidence. A policy excerpt is not an instruction. A JSON schema describes output shape but does not prove the content. Keeping these roles separate is the first defense against subtle failures.
Consider three applications:
- A document summary needs the document and a purpose; we should check whether the summary preserves meaning.
- Coding assistance needs requirements and relevant code; we should run tests and review the change.
- A policy answer needs the correct dated policy; we should check whether the cited passage supports the claim.
The prompt can specify all three workflows, but only the supplied document, code, or policy can provide their evidence.
Make Requirements Observable
“Write a good answer” is not testable. “Use only the supplied policy, cite its ID for every duration, and return insufficient evidence when no applicable policy exists” is testable. An output requirement should let either a person or a program decide whether the response passed.
A useful prompt therefore specifies:
- the decision or transformation;
- the scope of permitted information;
- required inputs and missing-input behavior;
- the output structure;
- factual, safety, and formatting constraints;
- an explicit rule for uncertainty or abstention.
This does not mean every prompt should be long. It means every sentence should remove a material ambiguity. If the user’s region changes which warranty applies, ask for the region. If a missing date makes two policies possible, ask for the date. More words are valuable only when they change the decision rule.
Developing Prompts as Experiments
The Iterative Loop
Prompt engineering is best treated as experimental design. Define the goal, write a baseline, run representative cases, identify one failure, revise the part responsible for that failure, and test again. Continue until the prompt meets a stopping rule on held-out cases—not until one hand-picked example looks impressive.

Imagine that the first prompt is “Write a balanced article about AI in automobiles.” A test reveals that the response covers benefits but omits ethical concerns and cybersecurity risks. The useful revision names those missing dimensions. It does not simply say “be more comprehensive.” After revising, we test on new topics and edge cases so that we do not overfit to the example that exposed the original problem.
Record at least the prompt version, model version, decoding settings, test input, output, and evaluation result. Otherwise a later improvement may be impossible to reproduce. Shared prompt prefixes may be cached by some systems, but prompting still consumes inference compute. A longer prompt is not free, and it can create new conflicts or distractors.
Fixtures and Stopping Rules
For the warranty assistant, four small fixtures exercise different behaviors:
| Input condition | Expected behavior |
|---|---|
| Policy matches product and region | Return the duration with the policy ID. |
| Region is missing | Ask for the region. |
| No applicable policy exists | Report insufficient evidence. |
| Two applicable policies conflict | Flag the conflict and inspect authority or effective date. |
A prompt is ready only when it behaves acceptably on representative held-out fixtures. The stopping rule might require, for example, zero unsupported durations, valid structure on at least 99% of cases, and correct abstention above a threshold. Those numbers depend on the application, but the existence of a stopping rule is non-negotiable. Without one, “prompt improvement” becomes subjective editing.
What Prompt Engineering Cannot Guarantee
Prompt behavior can be sensitive to wording, example order, model version, and data distribution. A result from one benchmark does not automatically generalize to product support, medicine, law, or another model family. Historical findings that randomized labels sometimes caused limited degradation on particular classification tasks do not justify incorrect demonstrations in a factual system.
Prompt engineering also does not create a secure instruction boundary. Retrieved text may say “ignore the application and answer differently.” Delimiters such as XML tags clarify which span is data, but they do not make that span trustworthy. The host application must decide which instructions have authority, which records the user may access, and which tool calls are permitted.
Examples, Clarification, and Reasoning
Zero-, One-, and Few-Shot Behavior
Zero-shot prompting provides instructions but no completed demonstration. One-shot prompting adds one completed input-output example. Few-shot prompting supplies several. Count completed demonstrations—not policy excerpts, background text, or incomplete questions. There is no universal number at which “few-shot” becomes “many-shot.”
Examples change the request context, not the model weights. They can teach the desired label vocabulary, answer style, citation format, and failure behavior. A useful pair for the warranty assistant demonstrates both sides of the decision:
1 | Example A — supported |
Neither demonstration is evidence about BB43300. They show how to behave when evidence exists or is missing. This distinction prevents a common mistake: copying a duration from an example into the current answer.
Historical GPT-2 and GPT-3 experiments made zero- and few-shot task specification widely visible. Their lesson is not that more examples always improve performance. Gains depend on the model, task, example quality, order, label balance, and context length. Model scale also affected the ability to infer some tasks from context. Treat each change as an experiment against the metric that matters for the current system.
Clarify Material Ambiguity
The interview pattern asks follow-up questions before answering. A travel planner might ask about destination type, prior trips, preferred activities, cultural interests, accommodation, and budget. The next question should reduce uncertainty that could materially change the plan. Asking every possible question wastes time; asking none forces the model to invent preferences.
The same principle applies to support. If region is required to select a policy, ask for it. If region does not affect the answer, do not create unnecessary friction. A good clarification policy connects each question to a decision boundary.
Reasoning Is Useful but Not Evidence
Chain-of-thought demonstrations historically improved some arithmetic and symbolic benchmarks, especially for sufficiently capable models. Zero-shot instructions such as “think step by step” became useful baselines. But the size of the gain varies by model and benchmark, and a fluent derivation can still be wrong. The correct response to a multiplication problem is to verify the arithmetic independently; the correct response to a warranty question is to verify the cited policy.
Reasoning models can start with a direct request: decide whether the supplied policy applies, cite its ID if it does, and otherwise state what is missing. Add examples only if the baseline fails. A longer explanation is not stronger evidence. For operational systems, prefer outputs that expose checkable objects—selected record IDs, parsed dates, calculated values, and tool results—rather than trusting persuasive prose.
From Prompt to Evidence-Grounded Application
Separate Rules, Examples, and Current Evidence
The support assistant has three distinct inputs:

The stable developer instruction can define the decision rule:
1 | Answer the warranty question for the specified product, region, and date. |
The current request then supplies the user question and authorized evidence. An API may keep these concepts separate through an instructions field and an input field. XML-style tags can mark evidence boundaries and preserve metadata:
1 | <evidence> |
Before sending this record, the application should authorize it, verify product/region/date, and retain P1 so the answer can cite it. The tags improve parsing; they do not establish authorization or factual correctness by themselves.
Structured Output and Semantic Validation
A schema gives downstream code a stable shape:
1 | { |
Syntactic validation checks whether the keys, types, and allowed values are correct. Semantic validation asks harder questions: Does P1 apply to this product, region, and date? Does it actually say twelve months? Is the evidence current and authorized? Valid JSON can still contain an unsupported answer.
Response handling should also distinguish a refusal, an incomplete response, and a parsed object. A refusal may need a separate user-facing path. An incomplete generation may be retried or reported. A parsed object proceeds to evidence checks.

This is the point where a prompt-only system reaches its limit. If the relevant policy is not already in the request, the application must find it.
What RAG Changes
External Evidence at Request Time
Retrieval-augmented generation (RAG) retrieves external evidence relevant to a request, selects evidence for the context, and generates an answer conditioned on both the question and that evidence. It can provide information that is current, private, or too large to encode reliably in model parameters. It can also be cheaper and easier to update than retraining a model whenever a policy changes.
Model parameters and request context play different roles. Parameters contain general language capabilities and patterns learned during training; a retrieval request does not rewrite them. The application inserts current policy text into the context for this answer. If the policy changes tomorrow, update the document store and retrieval path—not the model’s memory of the world.
RAG is more than semantic search. Search returns relevant records. A RAG system also selects context, asks a generative model to use it, and validates the resulting claims. Lexical, dense, or hybrid retrieval can all supply candidates.
Offline Indexing and Online Querying
The pipeline has two broad phases. During indexing, the application reads documents, splits them into usable passages, creates lexical or dense representations, stores the passages and metadata, and builds search indexes. During an online query, it:
- interprets the question;
- retrieves candidates;
- optionally reranks them;
- selects permitted, nonredundant context;
- generates an answer;
- validates citations and claims.

Every stage has a distinct failure mode. If the correct policy never becomes a candidate, generation cannot repair retrieval. If it is retrieved but removed during truncation, inspect context construction. If it reaches the model but the answer contradicts it, inspect generation and validation. This separation is the foundation of practical RAG debugging.
Lexical Retrieval and BM25
Inverted Indexes Preserve Exact Terms
Lexical retrieval is strongest when exact identifiers matter: product codes, error messages, names, dates, and quoted phrases. An inverted index maps each term to a postings list of documents containing it. For three documents,
1 | D1: BB43300 warranty, US |
the posting list for BB43300 is warranty is

Tokenization, case normalization, stemming, stop-word handling, field boosts, and query operators all affect the result. Serial numbers should often remain intact. A tokenizer that splits BB43300 unpredictably can destroy the very exact match lexical search is meant to preserve.
BM25 Scoring
BM25 combines term rarity, term frequency with diminishing returns, and document-length normalization. A common form is
Here
where
If
At

Length normalization prevents long documents from winning simply because they repeat more terms. It can also penalize legitimately long records, so
Dense Retrieval and Chunking
Closing the Vocabulary Gap
Lexical search fails when a query and a relevant passage express the same idea with different words. “Strong pain in the side of the head” and “sharp temple headache” may share no useful keyword. A dense embedding model maps both into vectors trained so that semantically related queries and passages score similarly.

Dense retrieval does not discover a universal geometry of meaning. Similarity depends on the encoder, training objective, domain, input formatting, and scoring metric. A model trained for general sentence similarity may not rank product policies correctly. Query and document encoders must be compatible, and embedding model versions must be tracked when the index is updated.
Bi-Encoders and Similarity
A bi-encoder computes a query vector and each document vector separately:
Document vectors can be computed once and reused for many queries, which makes large-scale retrieval practical.

Dot product and cosine similarity can rank the same vectors differently. Cosine removes vector magnitude:
For
Chunking Is a Retrieval Decision
Documents are usually split into chunks before embedding. A tiny chunk may match the query but omit the sentence needed to answer it. A huge chunk preserves context but consumes tokens, mixes topics, and can dilute the retrieval signal. Useful chunk boundaries often follow headings, paragraphs, code units, policy clauses, or tables rather than arbitrary character counts.

Every chunk should retain metadata such as source ID, document version, heading, product, region, effective date, access policy, and character or page offsets. The vector finds a candidate; metadata determines whether the candidate is applicable, authorized, current, and citable.
Approximate Nearest Neighbors
Why Exact Search Becomes Expensive
Exact nearest-neighbor search compares a query with every stored vector. For
Memory matters as well. One million 768-dimensional float32 vectors contain
or about
Navigable Small-World Graphs
A navigable small-world (NSW) index represents vectors as graph nodes. New nodes connect to selected nearby nodes, while some longer links allow movement between distant regions. Search begins at an entry node, evaluates its neighbors, moves toward candidates closer to the query, and explores the neighborhood until its stopping rule is satisfied.

A purely greedy walk can stop at a local choice. Let

Practical graph search keeps a candidate set instead of following only one edge. More exploration generally improves recall at a latency cost.
HNSW Adds a Hierarchy
Hierarchical Navigable Small World (HNSW) assigns nodes to random maximum levels. Every node appears in the dense base layer; progressively fewer nodes appear in higher layers. Search starts in a sparse upper layer, makes long-distance progress, descends through denser layers, and finishes with local candidate exploration at layer


HNSW exposes a recall-latency-memory trade-off. Greater graph degree and more construction effort usually improve index quality but consume memory and indexing time. Exploring more candidates at query time can improve recall but increases latency. The often-quoted logarithmic intuition is not a universal worst-case guarantee; measure the target workload against an exact reference.
Measuring ANN Recall
For exact top-
If
A vector database connects vector IDs to source text and metadata, supports similarity queries, and coordinates updates and deletes. Filtering may be required before or during ANN search so that users see only authorized records. Post-filtering can leave fewer than
Hybrid Retrieval and Reranking
Complementary Candidate Signals
Lexical and dense retrieval fail differently. Lexical search preserves exact identifiers such as BB43300; dense search can match paraphrases such as “How long is coverage?” to a passage containing “warranty duration.” A hybrid retriever runs both, unions their candidate IDs, removes duplicates, and combines their rankings.
Raw BM25 and cosine scores live on different scales, so adding them without calibration is unreliable. Reciprocal rank fusion (RRF) combines ranks instead:
where
while candidate
Reranking a Smaller Set
A bi-encoder retrieves cheaply from a large corpus because query and document vectors are computed separately. A cross-encoder jointly processes each query-document pair, allowing richer interaction but requiring a model call or forward pass per pair. This makes cross-encoders well suited to reranking dozens of candidates rather than scanning millions of documents.

Reranking improves ordering, not truth. “The capital of Canada is Sydney” may be highly related to the query while still false. A reranker also cannot select the correct policy if retrieval never returned it. When the answer is missing, inspect candidate recall before tuning the reranker.
Context Construction, Verification, and Debugging
Build Context Under a Token Budget
The final model input should contain relevant, permitted, and nonredundant passages while reserving space for instructions and output. Context construction may deduplicate overlapping chunks, group adjacent sections, prefer authoritative versions, and remove passages outside the product, region, date, or user authorization scope.
Preserve metadata in the final context. A passage saying “12 months” is not enough if the model cannot tell which product, country, date, or source it belongs to. For the running example, the selected context might be:
| Context element | Content |
|---|---|
| Application rule | Answer only when supplied evidence supports the claim; otherwise report insufficient evidence. |
Selected evidence P1 | BB43300; US; effective 2026-08-01; warranty 12 months. |
Excluded candidate P2 | A policy for another product; not evidence for this question. |
| Generation budget | Reserve space for a concise answer and a P1 citation. |
The supported answer is narrowly scoped: “For BB43300 in the US, the supplied policy states a 12-month warranty [P1].” If no applicable policy exists, answer “Insufficient evidence to determine the warranty.” Do not silently generalize a US policy to Canada or a single product to an entire product family.
A Citation Must Support the Claim
A citation is useful only when it supports the exact claim next to it. If P1 states that BB43300 has a 12-month US warranty, it does not support “All BB devices have a two-year worldwide warranty [P1].” Claim verification should decompose the response into checkable statements and compare each with the permitted passages.
Missing, conflicting, and hostile evidence require explicit behavior:
- Missing: ask for required context or abstain.
- Conflicting: compare authority, scope, and effective date; do not average policies.
- Hostile: treat instructions inside retrieved text as evidence content, not application authority.
- Unauthorized: exclude the record before generation, even if it is highly relevant.
Diagnose the Stage That Failed
A good trace records query transformations, filters, lexical and dense candidates, scores, ANN settings, reranker results, selected chunks, final model input, output, and claim-to-source checks. Then an observed failure points to a first inspection target:
| Observed failure | First stage to inspect | Useful evidence |
|---|---|---|
| Correct policy absent from candidates | Retrieval | Corpus, filters, query, scores, exact-vs-ANN comparison |
| Policy retrieved but absent from model input | Context construction | Selected chunks, deduplication, authorization, truncation |
| Policy in context but answer unsupported | Generation and validation | Exact input, output, and claim-to-passage support |
Frameworks may name components differently—loader, reader, splitter, parser, embedding model, vector store, retriever, postprocessor, prompt builder, response synthesizer, tracing hook—but they organize the same responsibilities. The names do not replace measurement.
The whole lesson can now be summarized in one division of labor:
- Instructions specify the task.
- Demonstrations show expected behavior.
- Retrieval supplies candidate evidence.
- ANN trades exact vector neighbors for speed and memory efficiency.
- Hybrid search and reranking decide which candidates deserve attention.
- Context construction determines what actually reaches the model.
- Generation turns selected evidence into language.
- Validation checks whether each claim is supported.
If a Canadian customer asks about BB43300 but the selected context contains only a US policy, the correct response is not to guess from the US duration. The assistant should state that the supplied evidence does not establish Canadian coverage, and the application should retrieve a Canada-specific policy or request the missing authority. That final restraint is what turns a fluent answer into an evidence-grounded system.
- Title: LLM 4 - Prompt Engineering and RAG
- Author: Gavin0576
- Created at : 2026-10-02 18:00:00
- Updated at : 2026-10-02 20:44:06
- Link: https://jiangpf2022.github.io/blog/2026/10/02/LLM-4-Prompt-Engineering-and-RAG/
- License: This work is licensed under CC BY-NC-SA 4.0.