suo
Guides

Measuring relevance with the eval CLI

Grade your own search against a golden set of queries, so a change to ranking is measured, not just eyeballed.

suo eval runs a set of judged queries against your index and reports standard relevance metrics, the same way suo's own engine is graded in CI before a ranking change ships.

Writing a golden set

A golden set is a list of queries with graded expected results, 0 (irrelevant) to 3 (exactly right):

[
	{
		"query": "reset my password",
		"judgments": [
			{ "objectID": "page-account-security", "grade": 3 },
			{ "objectID": "page-account-overview", "grade": 1 }
		]
	}
]

Anything not listed for a query is treated as ungraded and excluded from that query's score, not counted as irrelevant — so a golden set does not need to judge every record, only the ones that matter for that query.

Running it

npx @suo/cli eval --set golden-sets/docs.json --index docs
Query                         NDCG@10   MRR    Recall@10
reset my password             0.91      1.00   1.00
export my highlights          0.68      0.50   0.75

Overall                        0.83      0.79   0.89
Zero-result rate               1.2%
p95 latency                    41ms

Metrics

MetricAnswers
NDCG@10Are the best results near the top, weighted by how relevant each one is?
MRRHow far down the list is the first relevant result, on average?
Recall@10Of everything graded relevant, how much made it into the top 10?
Zero-result rateWhat share of queries returned nothing?
p50 / p95 latencyHow fast the queries came back

Comparing before and after a change

Run suo eval before and after a settings change — a new synonym, a custom ranking rule, a searchable tier reweighted — and compare the two reports directly. This is the same discipline the engine team uses internally: a paired comparison against the previous run, not a single number in isolation.

CI

Add suo eval --set golden-sets/docs.json --index docs --fail-under ndcg=0.75 to your CI pipeline to catch a regression in relevance before it reaches production, the same way a test suite catches a regression in code.

On this page