Use cases
LLM judging at production scale
“Is this answer supported by the context?” Cheap enough to run on every response, not a sample.
The problem with LLM-as-judge
Using a reasoning model to grade another model’s output is slow and costly, so teams sample. Sampling means most bad answers ship unchecked.
The Simo version
In a RAG pipeline, each response carries a support probability.
P(supported by context) → 0.97 ship
0.41 regenerate or abstainBelow the threshold the response regenerates or abstains. Ask more than one question in the same request: is it supported, does it answer the question, does it contain anything unsafe?
Related
See how Simo works and latency for why this fits on every response.
Frequently asked questions
Can I judge every response?
That is the design goal: a judgment takes milliseconds and ten questions cost little more than one, so it is cheap enough to run on all traffic.