evo-ai — Measuring the assistant
Traces, a scored corpus, and grounding metrics — the whole loop on one workstation.
A test suite tells you the code still runs. It cannot tell you whether an AI assistant answered better than it did last week. This is the evaluation loop built around evo-ai to answer that question with numbers — self-hosted tracing, a versioned question corpus, an LLM judge and grounding metrics — and an honest account of which of those numbers are allowed to decide anything.
The question a test suite cannot answer
evo-ai has a large conventional suite, and it is the wrong instrument for this. A unit test asserts that a function returns what it returned before. The thing that actually matters about a retrieval assistant — did it answer the question, from the records, without inventing anything — is a judgement, and it changes when nothing in the code changed at all: a model alias moves, a prompt is reworded, a retrieval threshold shifts by a tenth.
So the measurement has to be its own apparatus: a fixed set of questions, run against a deployed service, scored the same way every time, with each run kept so two runs can be compared. The three things that apparatus must produce are a verdict (did this answer pass), a reason (why), and a trace (what the service actually did to produce it). Without the third, a failing verdict is an accusation with no evidence.
The loop
Four pieces, each replaceable:
| Piece | Does |
|---|---|
| The corpus | Ninety-eight cases in a plain JSONL file — the question, deterministic checks, and the tags that say which suite it belongs to. Versioned in git like code, because it is. |
| The runner | evo.ragframework. Asks a deployed service each question through an adapter, applies the checks, calls a judge model for the ones a check cannot decide, and writes every answer to SQLite so a run can be re-read without asking again. |
| The observability platform — Langfuse | Self-hosted Langfuse. Receives the service's own spans, and mirrors the corpus as a dataset, each run as a dataset run, and every verdict as a score against the trace that produced it. |
| The metrics stage — DeepEval | DeepEval scoring each answer for grounding — faithfulness, relevancy, contextual precision and recall, and a rubric — against the chunks the answer actually cited. |
The seam that matters is the adapter. A case talks to an interface, not to evo-ai: there is an HTTP adapter that asks a running service, and a replay adapter that feeds recorded answers and recorded sources back through the identical scoring path. Replay is what makes it possible to ask "would today's metrics have caught last quarter's bug?" without a time machine — and it is what the whole exercise in the companion post rests on.
The tooling
Everything in the loop, named. All of it is open source except the hosted judge model, and every version shown is the one pinned — an evaluation stack whose own parts drift between runs measures its own drift.
The harness
| Tool | Role |
|---|---|
evo.ragframework (ragf) | The harness itself, built for this: the JSONL corpus, adapters for the service under test (HTTP and replay), deterministic checks, the runner, reports, and side-by-side comparison of any two runs. |
| httpx | The HTTP adapter's client to the service under test. |
| SQLite | Every answer, source, verdict and score from every run, so a run can be re-read, re-scored and compared without asking the service again. |
| pytest + GitHub Actions | The harness's own test suite, on every push. |
Judging and grounding metrics
| Tool | Role |
|---|---|
| DeepEval 4.2 | The grounding metrics: Faithfulness, Answer Relevancy, Contextual Precision, Contextual Recall, Hallucination, and G-Eval scored against each case's written rubric. |
| LiteLLM | One interface to whichever judge model is chosen, hosted or local, so changing the judge is a flag rather than a code change. |
| Gemini 2.5 Pro | The hosted judge and metric model — deliberately a different model from the one answering the questions. Its daily quota is a real budget line (see below). |
| Ollama running Llama 3.1 8B | A local metric model for exercising the pipeline offline. Not trusted as a judge: it returned zeros and reasoned about JSON formatting rather than the answer. |
Tracing and observability
| Tool | Role |
|---|---|
| Langfuse 4.50, self-hosted | Traces, datasets, dataset runs and scores in one place — web UI and ingestion worker as two containers. |
| ClickHouse 25.12 | Langfuse's columnar store for traces, observations and scores, and most of the reason the stack wants sixteen gigabytes. |
| PostgreSQL 17 | Langfuse's transactional data: users, projects, API keys, dataset definitions. |
| Redis 7 | Langfuse's ingestion queue and cache. |
| MinIO | S3-compatible object storage for raw events and media; the one other port, besides the UI, that is published. |
| OpenTelemetry | The tracing SDK inside evo-ai: the tracer provider, parent-based sampling, and W3C Trace Context propagation that lets a harness run and the service's spans become one trace. |
| OpenInference | openinference-instrumentation-llama-index, which emits LlamaIndex's own retriever, embedding and generation spans without hand-instrumenting each call. |
| Langfuse Python SDK v4 | OTLP export from evo-ai, and dataset, run and score publishing from the harness. |
Runtime and reproducibility
| Tool | Role |
|---|---|
| Docker Compose on Docker Desktop | Runs the six-container Langfuse stack and a production-shaped copy of evo-ai side by side on the workstation, so runs go against the same build that ships rather than a development server. |
| uv | evo-ai's lockfile, and the image installs from it frozen — so the service under evaluation is exactly the dependency set that was tested. It was not always: an unlocked rebuild once pulled a model-client release that broke every question. |
The system under test
evo-ai itself: FastAPI, LlamaIndex, Qdrant hybrid retrieval with fastembed, and generation through LiteLLM for hosted models or Ollama for local ones.
Self-hosted, and why that is the default
Langfuse runs as six containers on a workstation — its web UI and ingestion worker, plus PostgreSQL, ClickHouse, Redis and MinIO — under Docker Compose, from a file adapted from upstream and kept in its own repository.
It is not on the production box, and that was a sizing decision rather than a preference. Upstream's own guidance for this stack is four cores and sixteen gigabytes, most of it ClickHouse. The production box has two cores and under eight, and is already running dozens of containers for things customers actually use. Putting an analytics warehouse on it to measure a chat feature would have been the tail wagging the dog. So there is no public route, no edge vhost and no deploy target for it — a deliberate absence, revisited when a nightly run or a second reader needs one.
Adapting the upstream compose file meant four changes, and each one is held by a test in that repository:
- Pinned to an exact release, never a moving tag — an observability platform that silently upgrades itself is a measurement you cannot reproduce.
- Every secret is required with no default. Upstream ships working development values; a stack that boots with a default secret is a stack that can be left that way.
- Only two ports are published — the web UI and the object store the browser uploads to. Upstream also binds Postgres to the loopback address, which on this machine was already another project's development database.
- Telemetry off.
The point of the tests is not that the compose file is complicated. It is that every one of those choices is invisible in normal use and silently reversible on the next upstream merge.
What one question looks like
evo-ai is instrumented at the LlamaIndex layer, through OpenInference and OpenTelemetry, rather than at the model client — because it does not always go through one. The Ollama path bypasses LiteLLM entirely, and an instrumentation that only sees hosted calls would go blind in exactly the deployment mode the product is designed around. The whole thing sits behind a flag that is off by default, including in production.
A traced question arrives as a tree:
- a root span carrying the route the planner chose, whether the relevance gate refused, the top retrieval score, and the per-stage timings;
- a child span holding the chunks that were retrieved and their scores;
- beneath those, the framework's own retriever, embedding and generation spans.
The harness injects a W3C traceparent header when it asks, so the run and the service's own
spans are one trace rather than two systems describing the same event. The service joins that trace
only when its own sampling gate admits the caller — the harness can ask to be traced, it cannot insist.
The corpus, as data rather than a script
Each run mirrors the corpus into the platform as a dataset and the run itself as a dataset run, so a case has a history rather than a last-known result. Every case produces scores attached to its trace: the pass verdict, each named check, the judge's categorical verdict and its written reason, the latency, and any error.
The judge — Gemini 2.5 Pro, reached through LiteLLM — is deliberately a different model from the one under test, and its call is recorded as a generation span of our own — so the thing grading the answer is itself visible, costed and re-readable, rather than an opaque pass/fail.
One structural rule runs through the whole corpus, learned the expensive way on an earlier baseline: a case must assert a rule, never a row. Cases that asserted the contents of a seeded database passed until the database was reseeded and then failed for no product reason. Every one of those was retargeted onto the behaviour it existed to protect. The same trap recurs on every environment switch: on one recent run, four of seven judged failures were the local seed having no data of the kind the question asked about — the planner's SQL was correct and returned zero rows. A judged run has to be read against the seed before any failure becomes a bug report.
Grounding metrics
Checks and a judge tell you whether an answer was acceptable. They do not tell you whether it was grounded — whether every claim traces to a retrieved chunk. That is what the metrics stage adds, scoring each answer with DeepEval against the evidence it actually cited.
Measured across the regression cases that carry a written reference, at three trials each, medians:
| Metric | Median | What it is actually telling you |
|---|---|---|
| Faithfulness | 1.00 | Every claim supported by the rows. The one worth gating. |
| Answer relevancy | 1.00 | Nothing. It scores a refusal and a confidently wrong number identically — both are "relevant" to the question. |
| Contextual precision | 1.00 | Trivially true for an analytics answer: there is only ever one result node to rank. |
| Contextual recall | 0.50 | Measures the reference, not the answer. Half of a rule-shaped reference ("never reads as…") has no sentence in a SQL result that could support it. |
| Hallucination | 1.00 | Reads backwards from the name — in DeepEval’s current version it is the share of context the answer agrees with, so 1.00 means nothing was contradicted. |
That table is the useful output of the stage, and most of it is a warning. Three of those five metrics would be actively misleading as a gate, and only reading their reasons rather than their numbers reveals it. The stage therefore records each metric's own verdict rather than comparing a score to a threshold, because a metric whose direction you have assumed is worse than no metric.
One trial is not a measurement. Three trials of the same metric on the same answer came back 1.0, 1.0 and 0.2. The median held; a single run would have reported whichever one it drew. Anything recorded is a median of three.
What a number is allowed to decide
Nothing, yet — and that is the design, not an incomplete step. Thresholds are recorded, not enforced, for a stated settling period before any of them is allowed to fail a build.
The reason is that the metrics had to earn it. The bar set for the stage before it was built was explicit: it must be able to see two grounding defects that were already known, already fixed and already written up — on the old answers — and clear the fixed ones. If it could not see those, it was not worth enforcing.
On the first attempt it failed that bar, and the reason turned out to be a defect in evo-ai rather than in the metric. That story is the companion post; the short version is that the evidence the API returned for an analytics answer was the planner's own instructions rather than the rows, so the metric was grading against the wrong text and could not have caught anything. After that was fixed, the same replay scored the hallucinated answer 0.00 across three trials, naming the real figure in its reason.
A hosted judge is also a real budget line rather than a free assertion: one model alias exhausted its daily quota after a handful of judged corpus runs, and a nightly full-corpus run is comfortably above the free tier of the platform's hosted offering. That arithmetic is why the stack is self-hosted and why a nightly cadence is a decision with a cost attached rather than a default.
What this does not measure
In the interest of not overclaiming an evaluation harness, which would be a particularly circular failure:
- It measures a corpus, not the product. Ninety-eight questions are not the space of things people ask. The coverage suite — questions nobody had asked before — is the part that moves, and it currently passes well under half.
- A judge model is a model. It has been wrong in both directions here, including one apparent cross-tenant leak that was the judge misreading a workspace's own records as another's. Isolation held; the rubric was the thing that needed fixing.
- The numbers from two environments are not comparable. Different seed, different judge model, different build — a figure from one is not a regression against a figure from the other, and treating it as one manufactures bugs.
- Nothing here is automated yet. Runs are deliberate. A nightly cadence has a cost, a quota and a decision behind it, and is tracked as its own piece of work.
What it does buy is the thing that was missing: when the assistant gets worse, there is a number that says so, a reason attached to it, and a trace underneath showing what the service actually did.