RAG relevance

Check whether a retrieved chunk answers the query before calling the LLM, and decide when to retrieve more context, with the RAG Relevance template.

Goal: stop sending irrelevant chunks to your LLM. Score each retrieved chunk against the query, keep the useful ones and retrieve more when nothing answers the question.

Create the decision

Templates → RAG Relevance → Create decision:

PartContent
Statequery (string, required), chunk (string, required)
relevantprobability — does the chunk help answer the query?
relevancescore — none, partial, direct
retrieve_moreprobability — should the system retrieve more context before answering?
Policiesnone — every answer returns continue

Call it

curl -X POST https://api.dcision.io/v1/decisions/rag-relevance \
  -H "Authorization: Bearer $DCISION_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "state": {
      "query": "What is the refund window for annual plans?",
      "chunk": "Annual plans can be refunded within 30 days of purchase."
    }
  }'
200 OK
{
  "decision_id": "dec_Gt5Mv1Qz7Wp3Kx9Rc2Ln",
  "execution_id": "exec_Bw8Ls4Ty2Dq6Hn9Vk1Zf",
  "schema": "rag-relevance",
  "version": 1,
  "result": { "relevant": 0.96, "relevance": "direct", "retrieve_more": 0.08 },
  "confidence": { "relevant": 0.92, "relevance": 0.865, "retrieve_more": 0.84 },
  "scores": { "relevance": 2.91 },
  "action": "continue",
  "action_reason": { "type": "default" },
  "metrics": { "latency_ms": 334, "engine": "jev", "model": "jev-1.13.0", "estimated_cost_usd": 0.00001092, "input_tokens": 260, "output_tokens": 30 }
}

Add policies

The template has no rules, so the action is always continue. Two rules make the action meaningful — drop irrelevant chunks, ask for more context:

[
  { "field": "relevance", "on": "output", "operator": "eq", "value": "none", "action": "block" },
  { "field": "retrieve_more", "on": "output", "operator": "gte", "value": 0.7, "action": "fallback" }
]

Act on it

Score the retrieved chunks in parallel — each one is a decision — keep what passes, and only then call the LLM:

async function relevantChunks(query, chunks) {
  const decisions = await Promise.all(
    chunks.map(async (chunk) => {
      const response = await fetch("https://api.dcision.io/v1/decisions/rag-relevance", {
        method: "POST",
        headers: { Authorization: `Bearer ${process.env.DCISION_API_KEY}`, "Content-Type": "application/json" },
        body: JSON.stringify({ state: { query, chunk: chunk.text } }),
      });
      return response.ok ? { chunk, decision: await response.json() } : { chunk, decision: null };
    }),
  );

  const kept = decisions
    .filter(({ decision }) => !decision || decision.action !== "block") // keep on errors
    .sort((a, b) => (b.decision?.result.relevant ?? 0) - (a.decision?.result.relevant ?? 0));
  const needMore = decisions.every(({ decision }) => decision?.action === "fallback");
  return { kept: kept.map(({ chunk }) => chunk), needMore };
}

Tune it

  • Mind the volume: 10 chunks per question is 10 decisions and 10 requests. Size retrieval and your rate limit together, and score only the top results of your retriever.
  • Or score a fixed-size shortlist in one call: put the top passages in the state — { "query": …, "passages": [ … ] } — and define one probability question per position, such as passage_3: "Does `passages[3]` help answer `query`?" (up to 64 questions per decision). One request and one billable decision for the whole shortlist; keep the state within the token budget.
  • Turn off storeInput if chunks contain private documents.
  • Describe the corpus in context ("internal help center for a payroll product") so the engine can judge partial matches.
  • Use relevant for ranking and relevance for gating: the probability sorts chunks, the level drops the useless ones. For a single number, combine them in a composite.

On this page