Measuring relevance with the eval CLI
Grade your own search against a golden set of queries, so a change to ranking is measured, not just eyeballed.
suo eval runs a set of judged queries against your index and reports standard relevance metrics, the same way suo's own engine is graded in CI before a ranking change ships.
Writing a golden set
A golden set is a list of queries with graded expected results, 0 (irrelevant) to 3 (exactly right):
[
{
"query": "reset my password",
"judgments": [
{ "objectID": "page-account-security", "grade": 3 },
{ "objectID": "page-account-overview", "grade": 1 }
]
}
]Anything not listed for a query is treated as ungraded and excluded from that query's score, not counted as irrelevant — so a golden set does not need to judge every record, only the ones that matter for that query.
Running it
npx @suo/cli eval --set golden-sets/docs.json --index docsQuery NDCG@10 MRR Recall@10
reset my password 0.91 1.00 1.00
export my highlights 0.68 0.50 0.75
Overall 0.83 0.79 0.89
Zero-result rate 1.2%
p95 latency 41msMetrics
| Metric | Answers |
|---|---|
| NDCG@10 | Are the best results near the top, weighted by how relevant each one is? |
| MRR | How far down the list is the first relevant result, on average? |
| Recall@10 | Of everything graded relevant, how much made it into the top 10? |
| Zero-result rate | What share of queries returned nothing? |
| p50 / p95 latency | How fast the queries came back |
Comparing before and after a change
Run suo eval before and after a settings change — a new synonym, a custom ranking rule, a searchable tier reweighted — and compare the two reports directly. This is the same discipline the engine team uses internally: a paired comparison against the previous run, not a single number in isolation.
CI
Add suo eval --set golden-sets/docs.json --index docs --fail-under ndcg=0.75 to your CI pipeline to catch a regression in relevance before it reaches production, the same way a test suite catches a regression in code.