Guides
Everything about Simo
Simo is a judgment model, not a chat model. These guides cover how it works, why we built it, the two models, the published numbers, and where it fits.
Foundations
- How Simo worksHow a Simo answer is made: software shows it a situation, asks typed questions, and gets a calibrated probability for every option in milliseconds. No prose, no parsing.
- What is System 1.5?System 1 is fast intuition, System 2 is slow deliberation. Simo is System 1.5: a one-pass reflex with calibrated probabilities, plus exactly the deliberation that acting needs.
- Why we built SimoTeams answer in-loop judgment calls by prompting a chat model and parsing its essay: slow, expensive, brittle, unquantified. Why Temprl Labs built Simo instead.
- Calibrated probabilitiesWhen Simo says 0.8, it is right about 80% of the time. How calibrated probabilities turn “the model thinks so” into an act, escalate or ask-a-human threshold.
- Question typesThe five typed question types Simo answers: Yes/No, Choose one, Rate on a scale, Fill a value and Extract from the input, and what each one returns.
Models
- Simo modelsTwo Simo models, one API: Simo-1 for real-time judgment and Simo-1 Pro for maximum judgment quality. Compare latency, use and when to choose each.
- Simo-1Simo-1 is the default Simo model: 64 ms for one question and 104 ms for ten over a screenshot. Built for agent loops, pipelines and per-event processing.
- Simo-1 ProSimo-1 Pro is Simo’s highest-quality judgment model for high-stakes gating and hard calls: 96.75% on DecideBench, 189 ms for one question.
- LatencyPublished Simo latency over a 1280×720 screenshot: Simo-1 64 ms for one question, 104 ms for ten. Why reading is the cost and extra questions are nearly free.
- AccuracyPublished Simo-1 Pro scores: 96.75% on DecideBench and 81.8% on Banking77, with calibration error 0.047 and a comparison against other models.
Use cases
- All use casesAgent supervision, act-think-ask gating, ticket triage, visual QA, moderation, LLM judging, visual inspection, document reading and video events with Simo.
- Agent supervisionSupervise browser and computer-use agents with Simo: is the task done, what is the next action, which button, how risky. One call per step, in milliseconds.
- Act, think or askRoute decisions by calibrated confidence: act at 0.95, call the reasoning model at 0.60, ask a human below that. How Simo gates fast and slow AI systems.
- Ticket and inbox triageRoute a support email, score urgency and sentiment, and extract the order ID in one Simo request. One call replaces an intent, sentiment and extraction stack.
- Visual QA assertionsAssert on what a screen actually shows with Simo: “is the error dialog visible?” answered in milliseconds, surviving redesigns that break every CSS selector.
- Moderation and policy screeningScreen thousands of items an hour against each policy with Simo: a calibrated probability per policy, auto-action above the line, review queue near it.
- LLM judging at scaleCheck every RAG response with Simo: a calibrated probability that the answer is supported by its context, cheap enough to run on all traffic.
- Visual inspectionOne photo, one Simo call: extract the tracking ID, judge whether the parcel is damaged and where, and route damaged items to a claims queue.
- Document and form readingRead invoices and forms with Simo: vendor, invoice number and amount straight from the page, as typed values. An integer is always an integer.
- Video event detectionDetect events in video with Simo: a calibrated probability timeline, such as “person enters restricted area”, spiking when it happens.
Compare
- Simo vs reasoning modelsSimo answers directly with a calibrated probability; a reasoning model writes a few thousand words first. Compare speed, output, calibration and the best way to combine them.
- Simo vs prompting a chat modelPrompting a chat model and parsing its reply is slow, brittle and unquantified. Compare it with Simo’s typed questions and calibrated probabilities.
Reference
Stop parsing essays. Start reading probabilities.
Tell us what your software needs to judge.
Request API access