simoby Temprl Labs

Foundations

Calibrated probabilities: when Simo says 0.8, it is right about 80% of the time

Every decision Simo makes carries a calibrated probability. That turns “the model thinks so” into an engineering threshold: act above it, escalate to a human below it.

What calibrated means

A probability is calibrated when it matches how often the model is actually right. Take every call where Simo said 0.8: about 80% of them are correct. Take every call at 0.95: about 95% are correct.

Calibration describes Simo’s decision probabilities and self-check scores. It is not a measure of how fluent or sure a sentence sounds, and it is not a raw generation confidence.

Why you want it

An uncalibrated score is only a ranking. A calibrated one is a rate you can plan around: if you act at 0.95, you know roughly how often those automatic actions will be wrong, and you can decide whether that is acceptable for the cost of a mistake.

  • Above the act threshold (for example 0.95): the action executes on Simo’s answer alone.
  • In the middle: the item goes to a review queue, or to a reasoning model, the only time the expensive call is made.
  • Below the ask threshold (for example 0.60): it goes to a human, with the probabilities attached.

An example gate

Proposed actionSimo’s probabilityWhat happens
Press “Place order”0.98Executes
Route ticket to billing0.97Executes
Remove comment under policy 4.20.89Review queue
Approve refund of 40.000.55To a human
Close the account0.31To a human

Move the thresholds and the split moves with them. That tunable dial is the point: the cost of an error differs per product, and so should the threshold.

Measured calibration

On Banking77 (route a customer message to the right intent), the calibration error of Simo-1 Pro’s Choice probabilities is 0.047. See the accuracy page for the full scores.

A long tradition

Forecasters have been scored on stated probabilities since 1950, when Glenn Brier proposed grading weather forecasts that way (Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1–3). Simo brings the same discipline to software decisions.

Frequently asked questions

What does a calibrated probability mean?

That the stated probability matches the real hit rate. Of the calls Simo scores at 0.8, about 80% are right.

What threshold should I use?

It depends on the cost of an error. A common pattern is to act at 0.95, call a reasoning model at 0.60, and ask a human below that. Because the probabilities are calibrated, you can tune the line to your own risk.

What is Simo’s calibration error?

On Banking77, the expected calibration error of Simo-1 Pro’s Choice probabilities is 0.047.

Stop parsing essays. Start reading probabilities.

Tell us what your software needs to judge.

Request API access