How to Plan a RAG Pilot: Scope, Test Questions and Success Criteria
A RAG pilot should not be a large platform in miniature. With clear test questions, it should show whether the data, retrieval and answers genuinely work for a specific business task.

01
At first, a RAG pilot sounds fairly simple: ingest a few documents, add a search function, connect a language model and start asking questions. Technically, an initial prototype can indeed be built quickly.
The more important question is not whether an answer appears, but whether that answer is dependable enough for day-to-day work. Does the system find the correct source? Does it recognise which version is current? Does confidential information remain hidden from unauthorised users? And does it clearly state when the available data is insufficient?
That is exactly what a pilot is for. It is not a smaller version of the eventual full solution. It is a controlled test designed to support a sound decision.
02
A Pilot First Needs Clear Boundaries
“We want to chat with all of our company knowledge” is not a useful pilot scope. The data sources, user groups and expectations are too diverse. A clearly defined business task with a recognisable beginning and end is a better starting point.
A simple example: a service team should be able to answer questions about a specific product group using approved manuals. This already narrows down the user group, document set and type of questions considerably.
Five points should be clear at the outset:
- User Group: Who will actually use the system during the pilot?
- Task: Which specific activity should become easier or faster?
- Source Scope: Which documents and systems are included, and which are deliberately excluded?
- Access Rights: Who is allowed to see which information?
- Decision: What should the pilot tell us that we do not know today?
A recent Fraunhofer IAO study on RAG highlights its ability to make internal company information accessible without employees needing to know the exact filing structure. This is precisely why the selected source scope must be prepared in a professionally sound and transparent way.
03
Test Questions Come From Daily Work, Not the Presentation
A RAG system almost always looks convincing in a short demo. A few simple questions are asked, the relevant documents are known and the answers look good. That says little about how the system behaves when faced with real questions.
As a pragmatic working size for a manageable pilot, an initial set of roughly 25 to 40 questions is a useful starting point. This is not a universal standard. The mix is what matters:
- Straightforward Matches: The answer is clearly stated in a valid document.
- Distributed Information: The answer must be assembled from several sections or documents.
- Different Terminology: The question uses different wording from the source.
- Outdated Versions: Both old and current document versions exist.
- Missing Information: The correct response is to refrain from giving an unsupported answer.
- Access-Control Cases: The information exists but must not be shown to the person conducting the test.
Each question needs an expected key statement, the valid source and a responsible subject-matter expert. Otherwise, the final assessment remains stuck at “sounds good” or “I do not like it”.
04
Evaluate Retrieval and Answers Separately
In simplified terms, RAG consists of two steps. First, the system retrieves relevant passages. The language model then uses them to formulate an answer. A poor answer can therefore have two different causes: either the wrong content was retrieved or the correct content was used poorly.
This distinction matters. If retrieval returns an outdated work instruction, even the best prompt cannot fix it. If retrieval finds the correct passage but the answer invents an additional claim, the problem lies elsewhere.
Established evaluation approaches likewise distinguish between the quality of retrieved content and the quality of the final answer. A company pilot does not need to turn this into a set of complicated metrics. A transparent checklist is often sufficient at the outset.
05
A Simple Success Matrix for the RAG Pilot
| Evaluation Area | Key Question | The pilot passes if ... |
|---|---|---|
| Retrieval Quality | Does retrieval find the valid sources? | the relevant documents are retrieved reliably for the test questions. |
| Grounding | Does the answer stay grounded in the retrieved content? | claims can be verified through the displayed sources. |
| Completeness | Does the system answer every important part of the question? | no critical information from the valid sources is missing. |
| Uncertainty | Does the system recognise when the available information is insufficient? | it does not guess when data is insufficient and instead stops clearly or asks a follow-up question. |
| Permissions | Does restricted content remain protected? | existing access rules remain effective in both retrieval and answers. |
| Practical Value | Does the result help with the specific task? | the pilot group can genuinely work faster or more reliably with it. |
The criteria must be defined before testing begins. Defining them only after the first results makes it easy to adjust the benchmark to whatever the prototype happens to do well.
06
The Pilot in Four Work Packages
1. Define the Business Task and Responsibilities
We define the user group, task and desired outcome. At the same time, we determine who approves sources and who evaluates answers from a subject-matter perspective.
2. Prepare the Data Scope
The selected documents are reviewed for currency, duplicates, file formats, metadata and permissions. The entire organisation does not need to be cleaned up, but the pilot scope must be understandable.
3. Build the Prototype and Run the Test
The technical setup is connected to the prepared sources. The full question set is then run through the system. Errors are not merely counted; they are assigned to a cause: source, retrieval, preparation, answer or permission.
4. Make a Decision and Define the Roadmap
The result is not a polished demo, but a decision: continue, improve specific areas or stop. If the outcome is positive, the next data sources, user groups and processes are defined.
07
Clear Stop Criteria Are Part of the Pilot
A pilot may also show that RAG is not currently appropriate for the selected use case. That is not a failed project, but a valuable finding before a larger investment is made.
Stopping or fundamentally revising the pilot is appropriate if:
- there is no dependable source for the most important questions,
- no subject-matter owner is responsible for keeping content current and approving it,
- permissions cannot be represented correctly within the selected data scope,
- the pilot group sees no practical value despite good retrieval results,
- errors cannot be traced and improved systematically.
The European Data Protection Supervisor highlights risks associated with RAG, including sensitive data, inappropriate access and indirect manipulation. A pilot should therefore test these issues from the outset rather than treating them as later operational concerns.
08
Conclusion: The Pilot Should Deliver a Decision
A good RAG pilot does not prove that a chat interface can generate answers. It shows whether a clearly defined business task works dependably with the available data, access rights and responsibilities.
This requires a narrow scope, realistic test questions and success criteria agreed in advance. That is how an interesting demo becomes a sound basis for decision-making.
In the Trixner digital solutions RAG Pilot we therefore begin with the business task and the data. Only then do we choose the technical architecture and assess step by step whether expansion makes sense.
Next Step
Apply the Question to Your Own Company.
In the discovery call, we assess your specific situation and define a realistic next step.
