Confidence and probabilities

What confidence means for each question type on one 0–1 scale, how it affects minConfidence, weighted levels, the full probability distributions and near ties.

The engine doesn't just pick an answer: it returns a probability distribution for every question. Dcision turns it into a confidence per question — always — and can return the full distribution on request.

Calibrated probabilities

Jev is trained with RLCD — reinforcement learning for calibrated decisions — to return probabilities that mean what they say. Across many predictions of a well-calibrated model, outcomes given a probability of 0.2 happen about 20% of the time, and outcomes given 0.8 about 80% of the time. That is a property of groups of answers, not a guarantee about any single one — which is exactly what makes thresholds, escalation and auditing work.

Confidence

confidence has one number from 0 to 1 per question, keyed like result. It follows TypeSafe's definition, so every question type shares one scale: 1 when all the probability is on one answer, 0 when it is spread evenly or the answer is a coin flip.

Question typeConfidence is…Formula
choicehow far the chosen option stands above an even split between all n options, other included(p_max − 1/n) / (1 − 1/n)
scorehow concentrated the distribution is around the most likely level — probability on a neighboring level lowers it less than probability far away1 − spread / spread of an even distribution, floored at 0
probabilitythe distance from a coin flip|2p − 1|
  • For choice and score questions, Dcision uses the confidence the engine reports, computed with these formulas; it applies them itself only if the engine omits it.
  • Examples: a choice with five options and p_max = 0.88 has confidence 0.85. A probability of 0.9412 has confidence 0.8824; 0.03 has 0.94; 0.5 has 0.
  • For score questions, spread is the probability-weighted distance, in levels, from the most likely level. With three levels, (0, 0.5, 0.5) — torn between neighbors — has confidence 0.25, while (0.5, 0, 0.5) — torn between opposite ends — has confidence 0.

Use confidence in policies ("on": "confidence") or with a question's minConfidence to route uncertain answers to a person — see Confidence-gated routing.

Minimum confidence on probability questions

Because the confidence of a probability is |2p − 1|, a threshold on it flags a band of probabilities around 0.5: an answer is below a minimum confidence c when p falls strictly between (1 − c) / 2 and (1 + c) / 2.

minConfidenceFlags probabilities between
0.50.25 and 0.75
0.60.2 and 0.8
0.80.1 and 0.9
0.90.05 and 0.95

The same holds for rules with "on": "confidence": purchase_intent confidence lt 0.6 matches every p between 0.2 and 0.8. The editor shows the band next to the Minimum confidence slider of probability questions.

Coming from v0.1?

Until v0.2 (October 5, 2026), the confidence of a probability was max(p, 1 − p). The new value is lower for the same answer — p = 0.8 had confidence 0.8 and now has 0.6. To keep the behavior of an old threshold t, use 2t − 1: an old 0.8 becomes 0.6. Choice and score questions still use the engine's own confidence; when a provider omits it, Dcision now applies the formulas above instead of the top probability. Review minConfidence and confidence rules on probability questions in the Playground, then deploy a new version.

Weighted levels

For every score question, the response also has scores: the probability-weighted level, counted from 1. With the distribution low 0.02 · medium 0.21 · high 0.71 · critical 0.06, the most likely level is high (3), and the weighted level is 1 × 0.02 + 2 × 0.21 + 3 × 0.71 + 4 × 0.06 = 2.81.

{
  "result": { "priority": "high" },
  "confidence": { "priority": 0.69 },
  "scores": { "priority": 2.81 }
}
  • Use it in rules with "on": "score" to act between levels — priority ≥ 3.5 fires when the weight leans toward critical.
  • Composites use it, normalized to 0–1, as the value of a score term.
  • Use it for thresholds and ranking, not to reconstruct an exact quantity: score levels are categories, not a ruler.

Full distributions

Add include=probabilities to the query string:

curl -X POST "https://api.dcision.io/v1/decisions/lead-qualification?include=probabilities" \
  -H "Authorization: Bearer $DCISION_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "state": { "message": "We need pricing for 500 users and want to start next month.", "company_size": 500 } }'

The response gains a probabilities object, between action_reason and metrics:

{
  "result": { "purchase_intent": 0.9412, "priority": "high", "route": "sales" },
  "confidence": { "purchase_intent": 0.8824, "priority": 0.69, "route": 0.85 },
  "scores": { "priority": 2.81 },
  "composites": { "lead_score": 0.8414 },
  "action": "continue",
  "action_reason": { "type": "default" },
  "probabilities": {
    "purchase_intent": { "yes": 0.9412, "no": 0.0588 },
    "priority": { "low": 0.02, "medium": 0.21, "high": 0.71, "critical": 0.06 },
    "route": { "sales": 0.88, "sdr": 0.09, "nurture": 0.02, "spam": 0, "other": 0.01 }
  }
}
Question typeDistribution
choiceone entry per option, in option order, other last
scoreone entry per label, from lowest to highest
probability{ "yes": p, "no": 1 − p }

All values are rounded to 4 decimals. include accepts a comma-separated list; probabilities is the only value recognized today. Distributions don't change the price of a call.

The Playground always shows distributions, and executions store them when storeOutput is on.

Near ties

Two options are a near tie when the runner-up has more than 20% probability and trails the top option by less than 15 points — for example { "sales": 0.46, "sdr": 0.40 }. The Playground flags it:

Near tie: the top options are close. Make the option descriptions more distinct (add exclusions).

A near tie usually means two option descriptions overlap. Rewrite them as mutually exclusive criteria ("some intent, but no budget or timeline yet") and test again — the comparison with the previous run shows whether the answer moved.

Using distributions in your code

  • Second best. When the action is escalate, show the reviewer the top two options with their probabilities.
  • Your own measure. The distribution is all you need to compute another statistic — the top probability, or the ratio between the top two options — when it suits your decision better.
  • Calibrate thresholds. Compare the confidence of past runs — in Executions or through GET /v1/executions — with what really happened before choosing a minConfidence.

On this page