RAG relevance
Check whether a retrieved chunk answers the query before calling the LLM, and decide when to retrieve more context, with the RAG Relevance template.
Goal: stop sending irrelevant chunks to your LLM. Score each retrieved chunk against the query, keep the useful ones and retrieve more when nothing answers the question.
Create the decision
Templates → RAG Relevance → Create decision:
| Part | Content |
|---|---|
| State | query (string, required), chunk (string, required) |
relevant | probability — does the chunk help answer the query? |
relevance | score — none, partial, direct |
retrieve_more | probability — should the system retrieve more context before answering? |
| Policies | none — every answer returns continue |
Call it
curl -X POST https://api.dcision.io/v1/decisions/rag-relevance \
-H "Authorization: Bearer $DCISION_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"state": {
"query": "What is the refund window for annual plans?",
"chunk": "Annual plans can be refunded within 30 days of purchase."
}
}'{
"decision_id": "dec_Gt5Mv1Qz7Wp3Kx9Rc2Ln",
"execution_id": "exec_Bw8Ls4Ty2Dq6Hn9Vk1Zf",
"schema": "rag-relevance",
"version": 1,
"result": { "relevant": 0.96, "relevance": "direct", "retrieve_more": 0.08 },
"confidence": { "relevant": 0.92, "relevance": 0.865, "retrieve_more": 0.84 },
"scores": { "relevance": 2.91 },
"action": "continue",
"action_reason": { "type": "default" },
"metrics": { "latency_ms": 334, "engine": "jev", "model": "jev-1.13.0", "estimated_cost_usd": 0.00001092, "input_tokens": 260, "output_tokens": 30 }
}Add policies
The template has no rules, so the action is always continue. Two rules make the action meaningful — drop irrelevant chunks, ask for more context:
[
{ "field": "relevance", "on": "output", "operator": "eq", "value": "none", "action": "block" },
{ "field": "retrieve_more", "on": "output", "operator": "gte", "value": 0.7, "action": "fallback" }
]Act on it
Score the retrieved chunks in parallel — each one is a decision — keep what passes, and only then call the LLM:
async function relevantChunks(query, chunks) {
const decisions = await Promise.all(
chunks.map(async (chunk) => {
const response = await fetch("https://api.dcision.io/v1/decisions/rag-relevance", {
method: "POST",
headers: { Authorization: `Bearer ${process.env.DCISION_API_KEY}`, "Content-Type": "application/json" },
body: JSON.stringify({ state: { query, chunk: chunk.text } }),
});
return response.ok ? { chunk, decision: await response.json() } : { chunk, decision: null };
}),
);
const kept = decisions
.filter(({ decision }) => !decision || decision.action !== "block") // keep on errors
.sort((a, b) => (b.decision?.result.relevant ?? 0) - (a.decision?.result.relevant ?? 0));
const needMore = decisions.every(({ decision }) => decision?.action === "fallback");
return { kept: kept.map(({ chunk }) => chunk), needMore };
}Tune it
- Mind the volume: 10 chunks per question is 10 decisions and 10 requests. Size retrieval and your rate limit together, and score only the top results of your retriever.
- Or score a fixed-size shortlist in one call: put the top passages in the state —
{ "query": …, "passages": [ … ] }— and define one probability question per position, such aspassage_3:"Does `passages[3]` help answer `query`?"(up to 64 questions per decision). One request and one billable decision for the whole shortlist; keep the state within the token budget. - Turn off
storeInputif chunks contain private documents. - Describe the corpus in
context("internal help center for a payroll product") so the engine can judge partial matches. - Use
relevantfor ranking andrelevancefor gating: the probability sorts chunks, the level drops the useless ones. For a single number, combine them in a composite.
Agent routing
Let an AI agent pick its first tool and decide when a task really needs a large language model, with the Agent Routing template.
Writing good questions
Phrase instructions, options and criteria that Jev answers reliably — literal wording, math and dates in code, less indirection, a filtered state, option order and no text generation.