Overview and calibration
Read a decision's Overview tab — runs, success rate, latency, cost, actions, reasons and per-question statistics — and act on every calibration suggestion, with the thresholds behind it.
The Overview is the first tab of every decision. It answers three questions about the period you pick: is it working, what does it cost, and where does it fail or hesitate — and what should change? The editor is in the Editor tab next to it; new and duplicated decisions open straight in the editor.
Pick the period — 7, 30 (the default) or 90 days — and the traffic: All traffic, API or Playground. Everyone in the workspace can see the Overview, Viewers included. Days are counted in UTC.
Indicators
| Indicator | What it is |
|---|---|
| Runs | Every execution of the period — successful or not. The arrow compares it with the previous period of the same length (new when that one had none). |
| Success rate | The share of runs without an error, with the number of errors. It turns red from 2% of errors. |
| Latency p50 · p95 | The median and 95th-percentile latency of successful runs. |
| Input tokens / run | The average input tokens of successful runs — what drives the engine's cost and latency. |
| Engine cost (est.) | The sum of the runs' estimated engine cost, and how many versions ran in the period. |
Below them:
- Runs per day — runs and, dashed, errors.
- Final action — the share of each action callers received (successful runs and engine-error fallbacks).
- Why — which part of the decision chose the action: A policy rule matched, Answered "other", Below minimum confidence, Engine failed → fallback or No rule matched (continue) — and how many times each rule matched (Rules matched: #1 × 120).
- Errors — failed runs by error code; each code opens Executions filtered on errors. Errors are never billed.
- Traffic — API and Playground runs. Versions — runs per deployed version, and the draft for Playground runs.
Questions
One card per question, from a sample of the latest 2,000 successful runs of the period. Statistics need the answers, so they only cover runs stored with the decision's storeOutput on.
| Metric | What it shows | Highlighted from |
|---|---|---|
| Answers | The share of each option or level. Probability questions show the distribution of p(yes) in tenths, 0 to 1. | — |
| Confidence | A histogram of 20 bars of 0.05, from 0 to 1. Confident answers live on the right. | — |
| Avg confidence | The average confidence of the answers. | below 0.6 |
| Below minimum | The share of answers below the question's minConfidence — runs that get the fallback action unless a policy matched first. | 20% |
| Near ties (choice, score) | The share of answers whose top two options or levels were less than 0.15 apart, and the pair confused most often. | 15% |
| Undecided (probability) | The share of probabilities between 0.35 and 0.65 — neither yes nor no. | 30% |
| "other" (choice) | The share of answers where no option fit. | 10% |
| Avg p(yes) (probability) | The average probability of yes. | — |
Tune on each card opens the question in the editor.
Destinations
For decisions with destinations, a table counts, per destination and for the period:
| Column | Counts |
|---|---|
| Returned | Answers handed back in the response — functions, fixed replies, and LLM and agent answers. |
| Delivered | Deliveries that succeeded: webhooks, API requests, workflows and agent hand-offs. |
| Failed | Deliveries that failed, and LLM or agent calls that failed. |
| Pending | Deliveries still waiting for an attempt. |
Calibration suggestions
The Calibration suggestions card turns the numbers into advice. Each suggestion quotes the measurement behind it, so you can check it, and links to the field to change with Fix in the editor. They are sorted by severity — high, medium, low, info.
- Below 20 runs in the period, the card only says that the sample is small — no advice based on noise.
- Per-question advice also needs 20 answers of that question in the sample.
- Suggestions point at the draft — what you edit — or at the deployed version when the draft is invalid.
| Suggestion | Shown when | Severity | What to do |
|---|---|---|---|
| Per-question statistics are off | The decision doesn't store outputs. | info | Turn on Store output to get per-question advice. |
| Only 12 runs in this period | Fewer than 20 runs. | info | Run representative inputs in the Playground, or send real traffic. |
| 5% of runs failed (mostly ENGINE_TIMEOUT, 41) | 2% or more of runs failed — high from 10%. | medium, high | Per code: raise timeoutMs or answer with the fallback action (ENGINE_TIMEOUT); spread bursts or use your own key (ENGINE_RATE_LIMITED); turn on the engine-error fallback (ENGINE_UNAVAILABLE); fix what callers send (INVALID_STATE); otherwise read the errors in Executions. |
| 34% of runs end in the fallback action (escalate) | 30% or more of runs got the fallback action because of other or low confidence. | medium | Read the per-question advice before loosening thresholds. |
| "route" answered "other" in 14% of runs | A choice question answered other in 10% or more of its answers — high from 25%. | medium, high | Look at those executions; add an option for the common case or widen the descriptions. |
| "route" is a near tie in 22% of runs (mostly "sales" vs "sdr") | The top two answers of a choice or score question were less than 0.15 apart in 15% or more of its answers — high from 30%. | medium, high | Contrast the two descriptions and add exclusions ("not for existing customers"); Jev leans toward the first option, so put the more specific one first. For levels, describe each one with concrete criteria. |
| Average confidence on "priority" is 0.52 | The average confidence of a question is below 0.6. | medium | Ask one literal question, split compound ones, point at the state fields that matter with backticks, or add background in context. |
| minConfidence 0.8 sends 27% of "route" to escalate | 20% or more of a question's answers fall below its minConfidence. | medium | The card simulates the share below 0.5, 0.6, 0.7, 0.8 and 0.9 and names the highest threshold that keeps it at 10% or less. Lower the threshold only if answers near it are usually right; otherwise improve the question first. |
| Uncertain "route" answers can trigger blocking rules | The question has no minConfidence, the decision has a block rule, and some answers fall below 0.7. | low | Set a minConfidence so uncertain answers get the fallback action instead of being blocked. |
| "purchase_intent" is undecided (p between 0.35 and 0.65) in 31% of runs | A probability between 0.35 and 0.65 in 30% or more of its answers. | medium | Add counts as yes and counts as no criteria. |
| "route" answered "sales" in 93% of runs | One answer of a choice or score question in 90% or more of at least 50 answers. | low | Either the traffic is uniform or the question doesn't discriminate: sharpen the other options, or remove the question to save tokens. |
| "spam" never chosen in 412 runs | An option or level was never chosen in at least 200 answers (other aside). | low | Check its description, or remove it if that case doesn't happen. |
| Rule #3 (IF route eq spam THEN block) never matched in 1,250 runs | At least 100 runs and a rule never matched. | low | An earlier rule may always win, or the value is out of reach. |
| Rule #1 (…) matches 97% of runs | At least 100 runs and a rule matched 95% or more of them. | low | The rule may be broader than intended — later rules never get a chance. |
| 2,480 input tokens per decision on average | More than 2,000 input tokens per run on average. | low | Shorten instructions and context, merge questions, send smaller states. |
| p95 latency is 3.4 s | The p95 latency is above 3 s. | low | Latency grows with input tokens; on tight deadlines, answer with the fallback action on engine errors. |
| Destination "sales_crm" failed 12% of deliveries | At least 10 deliveries delivered or failed, and 5% or more failed — high from 25%. | medium, high | Open an execution to read the receiver's answer, fix the destination and resend. |
Calibrate, step by step
Start from the top suggestion. Open Executions filtered on the decision and read a few runs behind the number — their inputs, answers and distributions.
Change the draft in the editor: option descriptions, instructions, criteria, a threshold, a rule. Change one thing at a time.
Replay those inputs in the Playground — Re-run this input in the Playground from an execution — and compare with the previous run.
Deploy and watch the next days with the API traffic filter: the Overview includes every version of the period, and Versions shows how the runs split between them.
Thresholds such as minConfidence should follow the cost of a wrong answer, not only the fallback rate: a refund route deserves a higher bar than "send an article". See Handling low confidence and other and Writing good questions.
Calibration with outcomes
The Overview measures how sure and how consistent the engine is — not whether it was right. To compare confidence with real outcomes, export runs with GET /v1/executions and join them with what happened in your system.
Handling low confidence and other
Design decisions that admit uncertainty — the reserved other option, minimum confidence, confidence policies, the fallback action, near ties and the engine-error fallback.
Idempotent retries
A production-ready Dcision client — idempotency keys, timeouts, which errors to retry, exponential backoff with Retry-After — in JavaScript and Python.