What Building a RAG Eval Suite in Swift Taught Me About AI Engineering
WhattaRAG is my on-device RAG app, built entirely in Swift, and one of the most useful AI engineering lessons I've picked up building it came from its eval suite. To make retrieval and prompt changes provable instead of felt, I built an eval suite with 242 test cases. Its biggest win wasn't confirming a fix, it was killing three plausible ones and proving the real bug is a missing computation, not a wording problem. Here's what the evals measure, what they ruled out, and why seven of eight gate clauses still fail.