Composite scoring

Break a judgment into atomic scores and combine them with weights you control — composites computed by Dcision and usable in policies. Built on the Resume Screening template.

Break a complex judgment into independent dimensions, score each one separately, and combine them with weights you control — in the decision, not in a prompt. One broad question ("is this a good candidate?") hides several judgments behind one answer; atomic scores expose them so you can inspect, tune and combine them.

Benefits: cost, reliability and speed. TypeSafe's guide: Composite scoring.

Why it works

  • Each dimension is a small, literal question — the kind Jev answers best — and all of them come back from one call.
  • The math stays out of the model. Dcision computes the weighted combination exactly; Jev is good at judgment and weak at arithmetic.
  • You can see how a score was built. Every term is in the response, so when the ranking surprises you, you know which dimension — or which weight — to change.
  • Several rankings from the same answers. Two composites over the same four questions cost nothing extra: no new question, no new engine call.

When to use it

Ranking or qualifying with several criteria: lead scores, resume fit, vendor risk, content quality, the priority of a backlog. Whenever you would write "consider A, B and C" in one instruction, ask A, B and C separately and combine them.

How it maps to Dcision

In TypeSafe's guideIn Dcision
One score question per dimensionOne score (or probability, or choice) question per dimension
py = answers["python_depth"].score / 4Automatic: a score term is normalized to 0–1 as (weighted level − 1) / (levels − 1)
ic_score = 0.40 * py + 0.10 * lead + … in your codeA composite with those weights: Σ(weight × value) / Σ|weight|
Rank and filter in codeRead composites in the response, and use composites in rules to set the action

Build it

In the app

Open Templates and choose Resume Screening (composite scoring). Name the decision Resume Screening so its slug is resume-screening, then Create decision.

The State has one required string field, resume.

Four score questions, five levels each, one per dimension:

QuestionLevels, lowest to highest
python_depthnone, basic, working, strong, expert
team_leadershipnone, informal, tech lead, manager, manager of managers
system_designnone, basic, working, strong, expert
generalistnarrow, some, broad, very broad, full stack and ops

In the Composites card (Editor tab), Add composite twice:

  • senior_ic — Senior IC fit: depth and design first: 0.4 × python_depth + 0.1 × team_leadership + 0.4 × system_design + 0.1 × generalist.
  • eng_manager — Engineering Manager fit: leadership first: 0.15 × python_depth + 0.4 × team_leadership + 0.2 × system_design + 0.25 × generalist.

In Policies, the composites appear under Composites in the field list:

  1. IF senior_ic ≥ 0.7 — THEN continue.
  2. IF eng_manager ≥ 0.7 — THEN continue.
  3. IF senior_ic < 0.35, + AND eng_manager < 0.35 — THEN block.
  4. IF senior_ic < 0.7 — THEN escalate: the middle band goes to a recruiter.

Test a few resumes in the Playground — the Composites card shows both values — then Deploy v1.

The decision schema

{
  "stateSchema": {
    "kind": "object",
    "fields": [{ "key": "resume", "type": "string", "required": true, "description": "Resume text" }]
  },
  "questions": [
    { "key": "python_depth", "type": "score", "instructions": "How deep is the candidate's Python experience?", "scale": ["none", "basic", "working", "strong", "expert"] },
    { "key": "team_leadership", "type": "score", "instructions": "How much team leadership has the candidate shown?", "scale": ["none", "informal", "tech lead", "manager", "manager of managers"] },
    { "key": "system_design", "type": "score", "instructions": "How strong is the candidate's system design experience?", "scale": ["none", "basic", "working", "strong", "expert"] },
    { "key": "generalist", "type": "score", "instructions": "How broad is the candidate's experience across the stack?", "scale": ["narrow", "some", "broad", "very broad", "full stack and ops"] }
  ],
  "composites": [
    {
      "key": "senior_ic",
      "description": "Senior IC fit: depth and design first",
      "terms": [
        { "question": "python_depth", "weight": 0.4 },
        { "question": "team_leadership", "weight": 0.1 },
        { "question": "system_design", "weight": 0.4 },
        { "question": "generalist", "weight": 0.1 }
      ]
    },
    {
      "key": "eng_manager",
      "description": "Engineering Manager fit: leadership first",
      "terms": [
        { "question": "python_depth", "weight": 0.15 },
        { "question": "team_leadership", "weight": 0.4 },
        { "question": "system_design", "weight": 0.2 },
        { "question": "generalist", "weight": 0.25 }
      ]
    }
  ],
  "policies": [
    { "field": "senior_ic", "on": "output", "operator": "gte", "value": 0.7, "action": "continue" },
    { "field": "eng_manager", "on": "output", "operator": "gte", "value": 0.7, "action": "continue" },
    {
      "field": "senior_ic",
      "on": "output",
      "operator": "lt",
      "value": 0.35,
      "and": [{ "field": "eng_manager", "on": "output", "operator": "lt", "value": 0.35 }],
      "action": "block"
    },
    { "field": "senior_ic", "on": "output", "operator": "lt", "value": 0.7, "action": "escalate" }
  ]
}

Call it

curl -X POST https://api.dcision.io/v1/decisions/resume-screening \
  -H "Authorization: Bearer $DCISION_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: candidate-3307" \
  -d '{
    "state": {
      "resume": "8 years building Python services (Django, FastAPI). Designed the event pipeline for 20M users at a fintech. Tech lead of 4 engineers for 2 years. Some Terraform and on-call experience."
    }
  }'
200 OK
{
  "decision_id": "dec_Hn6Vx2Kq8Zt4Wm1Lp7Rc",
  "execution_id": "exec_Cs9Lw3Tq7Xm1Vn5Kd8Pz",
  "schema": "resume-screening",
  "version": 1,
  "result": {
    "python_depth": "strong",
    "team_leadership": "tech lead",
    "system_design": "strong",
    "generalist": "broad"
  },
  "confidence": {
    "python_depth": 0.625,
    "team_leadership": 0.8167,
    "system_design": 0.75,
    "generalist": 0.7083
  },
  "scores": {
    "python_depth": 4.35,
    "team_leadership": 3.02,
    "system_design": 4.1,
    "generalist": 2.85
  },
  "composites": { "senior_ic": 0.7418, "eng_manager": 0.5982 },
  "action": "continue",
  "action_reason": { "type": "policy", "rule": 0 },
  "metrics": {
    "latency_ms": 391,
    "engine": "jev",
    "model": "jev-1.13.0",
    "estimated_cost_usd": 0.0000189,
    "input_tokens": 450,
    "output_tokens": 47
  }
}

How senior_ic was built from the weighted levels in scores — five levels, so each term is (level − 1) / 4:

python_depth     (4.35 − 1) / 4 = 0.8375   × 0.4
team_leadership  (3.02 − 1) / 4 = 0.505    × 0.1
system_design    (4.10 − 1) / 4 = 0.775    × 0.4
generalist       (2.85 − 1) / 4 = 0.4625   × 0.1
senior_ic = 0.335 + 0.0505 + 0.31 + 0.04625 = 0.74175 → 0.7418   (the weights add up to 1)

senior_ic is 0.74, above 0.7: rule 0 matched and the candidate moves on as a Senior IC. Note that result holds the most likely level of each question, while composites use the weighted level — python_depth is strong, but with 40% on expert its weighted level is 4.35.

Act on it

Rank candidates by the composite of the role you are hiring for, and let the action pick the next step:

const screened = await Promise.all(resumes.map((resume) => decide("resume-screening", { resume: resume.text })));

const shortlist = screened
  .map((decision, index) => ({ candidate: resumes[index], decision }))
  .filter(({ decision }) => decision.action !== "block")
  .sort((a, b) => b.decision.composites.senior_ic - a.decision.composites.senior_ic)
  .slice(0, 10);

Here decide is a small wrapper around POST /v1/decisions/{slug} — see Idempotent retries for a production-ready one.

Tune it

  • Change weights, not instructions. If the top candidates don't match your expectations, adjust the weights and deploy a new version; the questions — and their calibration — stay the same.
  • Describe every level. Short labels work, but a rubric per level is clearer to the engine. A structured level keeps the short label in the API: { "label": "strong", "rubric": "primary language across several projects" }.
  • Add dimensions as questions. A new criterion is one more question and one more term — still one call.
  • Use negative weights for red flags, for example − 2 × job_hopping in a composite.
  • Threshold, don't interpolate. A weighted level of 4.35 means "between strong and expert, closer to strong" — good for ranking and thresholds, not for reconstructing an exact number of years.
  • Keep a person in the loop. Screening affects people: the middle band escalates, and even continue should lead to a human interview, not an automatic decision.

Related: Composites · Lead qualification

On this page