Response grading

Grade an LLM response against your rubric — accuracy, completeness and tone — into a 0–1 grade, and send low grades to review, with the Response Grading template.

Goal: know how good your assistant's answers are, response by response, and catch the bad ones before they cost you.

Create the decision

Templates → Response grading → Create decision:

PartContent
Statequestion, response (strings, required), reference (string — a known answer or source text)
ContextThe rubric — edit it to match your standards
accuracyscore — wrong, mostly wrong, mixed, mostly right, fully right
completenessscore — misses it, partial, most of it, complete
tonescore — poor, acceptable, good, excellent
unsafe_promiseprobability — does it promise a refund, price or deadline?
Compositegrade = 0.5 × accuracy + 0.3 × completeness + 0.2 × tone, 0 to 1
Policiesunsafe_promise ≥ 0.7 → escalate; grade < 0.6 → escalate
Destinationsquality_review — function flagForQualityReview on escalate

Call it

curl -X POST https://api.dcision.io/v1/decisions/response-grading \
  -H "Authorization: Bearer $DCISION_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "state": {
      "question": "Can I get a refund on my annual plan after 2 months?",
      "response": "Yes, of course! We'"'"'ll refund you in full anytime.",
      "reference": "Annual plans can be refunded within 30 days of purchase."
    }
  }'

The state's response is "Yes, of course! We'll refund you in full anytime.".

200 OK (abridged)
{
  "schema": "response-grading",
  "result": { "accuracy": "wrong", "completeness": "partial", "tone": "good", "unsafe_promise": 0.97 },
  "composites": { "grade": 0.17 },
  "action": "escalate",
  "destinations": [{ "key": "quality_review", "type": "function", "function": "flagForQualityReview", "params": { "grade": 0.17 } }]
}

Act on it

Grade after the reply is generated — before sending, to block it, or afterwards, to track quality:

const d = await dcisionDecide("response-grading", { question, response, reference });
const outOfTen = Math.round(d.composites.grade * 100) / 10; // 0.17 → 1.7 / 10
if (d.action === "escalate") await flagForQualityReview({ question, response, grade: d.composites.grade });

Tune it

  • Write your rubric in context — it is versioned with the decision, so a grade always points to the rubric that produced it.
  • Change the weights of grade to match what matters most for you.
  • Executions keep inputs and outputs by default: the Executions page and the Overview show grades over time.

On this page