AI & LLMs
Evaluate RAG Before Tuning the Prompt
Build a small retrieval evaluation set and measure context quality before spending days rewriting prompts or changing models.
When a retrieval-augmented assistant gives a weak answer, teams often edit the prompt first. The model cannot answer from evidence it never received, so retrieval quality should be measured before generation style.
This guide focuses on the decisions that survive contact with production: clear boundaries, observable behavior, and a feedback loop that reveals when an assumption is wrong.
The problem worth solving
A demo can look convincing while failing common queries. Without labeled questions and expected source documents, changes to chunk size, embeddings, reranking, or prompts become guesswork.
The useful move is to make the hidden constraint explicit. Write down what must stay correct, what can be delayed, and how the system should behave when a dependency fails. That turns a vague idea into something a team can test.
A practical implementation
Collect real questions, label the documents that contain the answer, and calculate recall at a small cutoff. Log retrieved document IDs alongside the final response so failures can be separated into retrieval and generation errors.
function recallAtK(expected: Set<string>, actual: string[], k = 5) {
const hits = actual.slice(0, k).filter((id) => expected.has(id))
return expected.size === 0 ? 1 : hits.length / expected.size
}
The example is intentionally small. In a real project, add structured logs, metrics around the failure path, and tests for retries or partial results. Keep the interface narrow so the implementation can change without forcing every caller to change too.
What to measure
Measure the outcome rather than activity. For software, that may be latency, error rate, queue depth, or recovery time. For product work, it may be activation, retention, or the number of useful conversations. Review the signal on a regular cadence and record what changed.
Takeaway
Treat retrieval as its own product surface. A small, representative evaluation set creates a faster improvement loop than prompt changes made by feel.
Start with the smallest version that can teach you something, make its behavior visible, and improve it from evidence. That rhythm is more dependable than trying to design the final answer in one pass.
Further reading
Explore more AI & LLMs articles from this journal.
Tushar Sharma