Overview and calibration

Read a decision's Overview tab — runs, success rate, latency, cost, actions, reasons and per-question statistics — and act on every calibration suggestion, with the thresholds behind it.

The Overview is the first tab of every decision. It answers three questions about the period you pick: is it working, what does it cost, and where does it fail or hesitate — and what should change? The editor is in the Editor tab next to it; new and duplicated decisions open straight in the editor.

Pick the period — 7, 30 (the default) or 90 days — and the traffic: All traffic, API or Playground. Everyone in the workspace can see the Overview, Viewers included. Days are counted in UTC.

Indicators

IndicatorWhat it is
RunsEvery execution of the period — successful or not. The arrow compares it with the previous period of the same length (new when that one had none).
Success rateThe share of runs without an error, with the number of errors. It turns red from 2% of errors.
Latency p50 · p95The median and 95th-percentile latency of successful runs.
Input tokens / runThe average input tokens of successful runs — what drives the engine's cost and latency.
Engine cost (est.)The sum of the runs' estimated engine cost, and how many versions ran in the period.

Below them:

  • Runs per day — runs and, dashed, errors.
  • Final action — the share of each action callers received (successful runs and engine-error fallbacks).
  • Why — which part of the decision chose the action: A policy rule matched, Answered "other", Below minimum confidence, Engine failed → fallback or No rule matched (continue) — and how many times each rule matched (Rules matched: #1 × 120).
  • Errors — failed runs by error code; each code opens Executions filtered on errors. Errors are never billed.
  • Traffic — API and Playground runs. Versions — runs per deployed version, and the draft for Playground runs.

Questions

One card per question, from a sample of the latest 2,000 successful runs of the period. Statistics need the answers, so they only cover runs stored with the decision's storeOutput on.

MetricWhat it showsHighlighted from
AnswersThe share of each option or level. Probability questions show the distribution of p(yes) in tenths, 0 to 1.—
ConfidenceA histogram of 20 bars of 0.05, from 0 to 1. Confident answers live on the right.—
Avg confidenceThe average confidence of the answers.below 0.6
Below minimumThe share of answers below the question's minConfidence — runs that get the fallback action unless a policy matched first.20%
Near ties (choice, score)The share of answers whose top two options or levels were less than 0.15 apart, and the pair confused most often.15%
Undecided (probability)The share of probabilities between 0.35 and 0.65 — neither yes nor no.30%
"other" (choice)The share of answers where no option fit.10%
Avg p(yes) (probability)The average probability of yes.—

Tune on each card opens the question in the editor.

Destinations

For decisions with destinations, a table counts, per destination and for the period:

ColumnCounts
ReturnedAnswers handed back in the response — functions, fixed replies, and LLM and agent answers.
DeliveredDeliveries that succeeded: webhooks, API requests, workflows and agent hand-offs.
FailedDeliveries that failed, and LLM or agent calls that failed.
PendingDeliveries still waiting for an attempt.

Calibration suggestions

The Calibration suggestions card turns the numbers into advice. Each suggestion quotes the measurement behind it, so you can check it, and links to the field to change with Fix in the editor. They are sorted by severity — high, medium, low, info.

  • Below 20 runs in the period, the card only says that the sample is small — no advice based on noise.
  • Per-question advice also needs 20 answers of that question in the sample.
  • Suggestions point at the draft — what you edit — or at the deployed version when the draft is invalid.
SuggestionShown whenSeverityWhat to do
Per-question statistics are offThe decision doesn't store outputs.infoTurn on Store output to get per-question advice.
Only 12 runs in this periodFewer than 20 runs.infoRun representative inputs in the Playground, or send real traffic.
5% of runs failed (mostly ENGINE_TIMEOUT, 41)2% or more of runs failed — high from 10%.medium, highPer code: raise timeoutMs or answer with the fallback action (ENGINE_TIMEOUT); spread bursts or use your own key (ENGINE_RATE_LIMITED); turn on the engine-error fallback (ENGINE_UNAVAILABLE); fix what callers send (INVALID_STATE); otherwise read the errors in Executions.
34% of runs end in the fallback action (escalate)30% or more of runs got the fallback action because of other or low confidence.mediumRead the per-question advice before loosening thresholds.
"route" answered "other" in 14% of runsA choice question answered other in 10% or more of its answers — high from 25%.medium, highLook at those executions; add an option for the common case or widen the descriptions.
"route" is a near tie in 22% of runs (mostly "sales" vs "sdr")The top two answers of a choice or score question were less than 0.15 apart in 15% or more of its answers — high from 30%.medium, highContrast the two descriptions and add exclusions ("not for existing customers"); Jev leans toward the first option, so put the more specific one first. For levels, describe each one with concrete criteria.
Average confidence on "priority" is 0.52The average confidence of a question is below 0.6.mediumAsk one literal question, split compound ones, point at the state fields that matter with backticks, or add background in context.
minConfidence 0.8 sends 27% of "route" to escalate20% or more of a question's answers fall below its minConfidence.mediumThe card simulates the share below 0.5, 0.6, 0.7, 0.8 and 0.9 and names the highest threshold that keeps it at 10% or less. Lower the threshold only if answers near it are usually right; otherwise improve the question first.
Uncertain "route" answers can trigger blocking rulesThe question has no minConfidence, the decision has a block rule, and some answers fall below 0.7.lowSet a minConfidence so uncertain answers get the fallback action instead of being blocked.
"purchase_intent" is undecided (p between 0.35 and 0.65) in 31% of runsA probability between 0.35 and 0.65 in 30% or more of its answers.mediumAdd counts as yes and counts as no criteria.
"route" answered "sales" in 93% of runsOne answer of a choice or score question in 90% or more of at least 50 answers.lowEither the traffic is uniform or the question doesn't discriminate: sharpen the other options, or remove the question to save tokens.
"spam" never chosen in 412 runsAn option or level was never chosen in at least 200 answers (other aside).lowCheck its description, or remove it if that case doesn't happen.
Rule #3 (IF route eq spam THEN block) never matched in 1,250 runsAt least 100 runs and a rule never matched.lowAn earlier rule may always win, or the value is out of reach.
Rule #1 (…) matches 97% of runsAt least 100 runs and a rule matched 95% or more of them.lowThe rule may be broader than intended — later rules never get a chance.
2,480 input tokens per decision on averageMore than 2,000 input tokens per run on average.lowShorten instructions and context, merge questions, send smaller states.
p95 latency is 3.4 sThe p95 latency is above 3 s.lowLatency grows with input tokens; on tight deadlines, answer with the fallback action on engine errors.
Destination "sales_crm" failed 12% of deliveriesAt least 10 deliveries delivered or failed, and 5% or more failed — high from 25%.medium, highOpen an execution to read the receiver's answer, fix the destination and resend.

Calibrate, step by step

Start from the top suggestion. Open Executions filtered on the decision and read a few runs behind the number — their inputs, answers and distributions.

Change the draft in the editor: option descriptions, instructions, criteria, a threshold, a rule. Change one thing at a time.

Replay those inputs in the Playground — Re-run this input in the Playground from an execution — and compare with the previous run.

Deploy and watch the next days with the API traffic filter: the Overview includes every version of the period, and Versions shows how the runs split between them.

Thresholds such as minConfidence should follow the cost of a wrong answer, not only the fallback rate: a refund route deserves a higher bar than "send an article". See Handling low confidence and other and Writing good questions.

Calibration with outcomes

The Overview measures how sure and how consistent the engine is — not whether it was right. To compare confidence with real outcomes, export runs with GET /v1/executions and join them with what happened in your system.

On this page