What Building a RAG Eval Suite in Swift Taught Me About AI Engineering
I was doing device QA on WhattaRAG, my on-device RAG app, and I asked it a multiple-choice-shaped question against a document that didn’t contain the answer. I handed it all three plausible options in the question itself. It answered “Date.”
“Date” isn’t in the document. It isn’t in the question either. The model just produced it, confidently, as if it had read it somewhere.
That was the moment “just tweak the prompt” stopped being credible to me. If the model can invent a word that appears nowhere in its inputs, no amount of prompt wording is going to fix that on its own. I needed a way to tell the difference between “this change helped” and “this change felt like it helped,” so I built an eval suite. Its job was to turn retrieval and prompt changes into something provable instead of something felt.
Building WhattaRAG has been teaching me a lot of AI engineering, and I’m enjoying it more than I expected. This post is about that eval suite: what it measures, what it ruled out, and the state of things as of the last baseline, where seven of eight gate clauses still fail.
Why build it all in Swift
WhattaRAG is 100% native Swift, 4.4 MB on the App Store, with no bundled model. Generation runs on Apple’s on-device Foundation Models; embeddings come from Apple’s NLContextualEmbedding, 512 dimensions per chunk.
There’s no Python anywhere in this stack, and no LangChain or anything like it. Every piece of the pipeline is something I wrote by hand: a chunker, hybrid search that fuses keyword and vector results with reciprocal rank fusion, cosine ranking done in Swift with Accelerate and vDSP over vectors stored in SQLite, a prompt builder that enforces a context budget, and the eval harness itself. Later I pulled the pipeline out into its own Swift package, which is what lets the eval suite run with swift test on a Mac, against the real Apple model, with no simulator and no device in the loop.
That’s not a stance against any framework. It’s a consequence of building for one platform, in one language, by hand. When you write each piece yourself, you can’t hide behind a framework’s default behavior, because there isn’t one. And when a number in the eval suite moves, up or down, you already know which piece moved it, because you’re the one who wrote it. That traceability is most of why the eval suite is worth trusting: it’s measuring a pipeline where every component’s behavior was my choice, not a library’s.
Why on-device makes this harder
A few things about running entirely on-device make the fabrication problem worse than it would be with a server-backed model.
First, there’s no server to fall back on for heavier reasoning or a second-pass check: whatever runs, runs on the phone, within the phone’s budget.
Second, each request gets a fresh LanguageModelSession. Apple’s guidance (TN3193) is to start a new session per request rather than keep one alive, so there’s no persistent context the model can lean on across turns to notice it’s repeating a guess.
Third, and this is the one that matters most: when a document fits inside the model’s context budget, WhattaRAG skips retrieval and ranking entirely and hands the whole document to the model. I call this the single-pass bypass, and it’s not an edge case: it happens on 78% of turns. Once the whole document is just sitting in the prompt, there’s no ranked, scored set of chunks to compare the answer against. No question ever looks “unsupported,” because nothing was filtered out in the first place. The pipeline literally has no signal that says “this might not be in there.”
What the ruler measures
I split the suite into two layers that get treated very differently.
Retrieval recall@K, the standard metric for retrieval quality, is deterministic. It only needs the embedding model, not the language model, and it reproduces byte-identically across runs: the same run, five times, gives the same number every time. Because it’s reproducible, it’s the layer I treat as a hard, enforced gate.
Answer quality is generation-dependent. The same prompt, run twice on-device, can produce different wording, different completeness, sometimes a different verdict on whether it fabricated. So answer quality is measured and reported, but it doesn’t gate anything by itself. You can’t hold a build to a bar that moves on its own between runs.
The suite is a golden dataset of 242 test cases. 194 are a working set I tune against, and 48 are a held-out set I never touch, which is the standard guard against overfitting to your own tests. The held-out set exists to catch “teaching to the test,” a change that looks great on the cases I’m actively iterating on but does nothing (or worse) on cases I never looked at while making the change. That slice used to be 17 cases, which sounds like plenty until you do the math: at 17 cases, one single case is worth about 5.9 points, which blew past the 5-point tolerance the gate itself allows. A held-out set that coarse can’t tell a real regression from rounding, so I resized it to 48.
The cases are spread across 7 corpora. Three (a coffee-shop handbook, an e-ink notebook manual, a bookmark-manager doc set) are ones I wrote myself, specifically so nothing in them could be sitting in the model’s training data: a correct answer there can only have come from retrieval, not memory. Two are public-domain classics, Adam Smith’s Wealth of Nations and Shelley’s Frankenstein, that I picked deliberately because the model has almost certainly memorized them, which lets me test whether it’s citing the document in front of it or just reciting what it already knows. One is built around multi-chunk sections rather than a register, to test retrieval across chunk boundaries. And one has zero markdown headings on purpose, to exercise the plain-text chunker path that the other six corpora never touch.
The gate clauses
There are eight named clauses.
- No Weak Corpus: retrieval recall@K per individual corpus, needs to clear 95% (90% for the two memorized corpora).
- Finds The Evidence: retrieval recall@K across the whole working set, needs 95%.
- Doesn’t Invent: on questions where the document has no answer, the answer can’t assert anything the document doesn’t say. In other words, no hallucination. Required 100%, working and held-out. Across recorded runs it’s actually landed in the 31–50% range.
- Says It Doesn’t Know: when the model doesn’t know, it has to decline and say the document doesn’t cover it, not just go quiet.
- Stays In The Document: refusal rate on questions that are genuinely outside the document. Required 100%.
- Document Over Memory: on the two memorized corpora, when a fact in the text is deliberately altered, the model has to follow the altered document over what it “remembers,” counted only where the altered span was actually retrieved.
- No Weak Category: no single question category (factoid, multi-hop, aggregation, and so on) can fall below an 80% pass rate.
- Holds On Unseen: the held-out slice can’t trail the working set by more than 5 points, the overfitting check.
As of the last five-run baseline, seven of these eight fail. The only one that has ever passed is Stays In The Document.
Three things the suite killed
This is the part that made building the suite worth it. It didn’t just confirm improvements, it killed three plausible fixes I would otherwise have shipped.
Query rewriting. I had an on-device rewrite step that reworded the user’s question before retrieval, meant to help with badly-phrased queries. With rewriting toggled on versus off, the miss lists were character-for-character identical, including on the two cases it was specifically supposed to help. It cost about 719 ms per query for zero measured benefit. Deleted.
Retrieval-confidence abstain. This was modeled on how a competing app handles the same problem: refuse to answer if the best-matching chunk is farther than some cosine-distance threshold. I measured the actual distance distribution, and it was inverted. Questions the model should have declined (no answer in the document) clustered closer, more confident-looking, than questions it should have answered. The AUC (area under the ROC curve, in plain terms, how well a threshold on this signal separates the two cases, where 0.50 means no better than a coin flip and 1.0 means perfect) came out to 0.449. Below 0.50. This signal is actively backwards for this data. Shipped off.
Post-generation grounding check. This one checks, after the answer is generated, whether it asserts something the prompt never actually supplied. On the numbers that made it to a real enforcement decision, precision was 2 of 4: of the answers it would have withheld, only half were actually fabrications, the other half were correct answers it would have thrown away. Enforcing it would destroy one good answer for every real fabrication it caught. Not gate-ready.
There’s a smaller, more uncomfortable result from an earlier round of testing, back in August. An explicit “don’t guess” instruction added to the prompt made grounding worse, not better: 75/77 down to 74/77.
More recently, a separate “write the answer yourself, in at least one complete sentence” directive turned out to be a regression on its own, 43.8% fabrication versus 29.2% without it, and it was only caught because I’d added a dedicated refusal-turn control to compare against (25.0% with the refusal branch stated first). Without that control, the regression would have hidden inside normal run-to-run noise.
What did move
Not everything failed. Header consolidation, collapsing the refusal instruction and the quoting rule out of the system header and into a directive that sits right next to the question, moved fabrication from 29.2% down to 12.5%. Good news, except the same change pushed quote-dumping (pasting document text back verbatim instead of answering) up from 13% to 24.2%. Fixing one failure mode made a different one worse.
I then ran a separate follow-up experiment, restoring just a single sentence about quoting, with its own before/after runs: dumping went from 28.0% down to 8.7%. I only caught the original trade-off in the first place because I was scoring fabrication, quote-dumping, and content-free answers (a bare citation marker with no actual content) all at once, on the same runs. If I’d only tracked fabrication, I’d have shipped the first version and called it a win.
What I learned about noise
On-device generation isn’t deterministic, and the suite forced me to learn its noise floor instead of guessing at it. A single-run comparison moves by about ±4 cases just from run-to-run variance in generation; on the 102-case working set that widens to about ±6. The multi-turn conversational path is noisier still, roughly ±10 points run to run.
I didn’t run anything as formal as a test for statistical significance to bound these numbers. What I have is a noise band, not a p-value: a real number below which “it got better” isn’t a claim I’m allowed to make. A fix I logged on 2026-09-07 moved fabrication from 70.8% to 64.6%. That’s inside the ~10-point band. I wrote it up as not a fix, on purpose, even though the number on the page looked like progress.
Where it stands
Gate status hasn’t moved: seven of eight clauses fail, and Stays In The Document remains the only one that’s ever passed. Both proxies I tried for “does the model know it doesn’t know,” the distance-based abstain and the post-generation grounding check, failed for the same underlying reason. Neither one actually encodes whether the document addresses the question at all; both are trying to infer that from a side effect after the fact.
The two candidates I haven’t built yet are more expensive on purpose: a cheap pre-answer pass that explicitly asks whether the document covers the question before generating anything, or a retrieval change that ranks without dropping, so “nothing here matches” becomes something the prompt can say rather than something the model has to infer.
Building the suite didn’t get me a passing gate. What it got me was honesty about which fixes are real, and it turned three good-sounding ideas I would have shipped into things I could point at and say, measured, doesn’t work, here’s the number. That’s a smaller win than “fixed it,” but it’s the one that’s actually true right now.
If you want to try WhattaRAG yourself, it’s on the App Store: WhattaRAG.
Thanks for reading!
← All posts