Writing good questions

Phrase instructions, options and criteria that Jev answers reliably — literal wording, math and dates in code, less indirection, a filtered state, option order and no text generation.

Jev is fast, calibrated and good at common-sense judgment — and it reads your question literally. Most accuracy problems come from how a question is written, not from the engine. These rules follow TypeSafe's list of known jev-1.13 limitations and its guide to structured questions. Many of these limitations should ease in later Jev versions.

1. Say exactly what you mean

Jev answers the question you wrote, not the one you meant. Scoping words, negations and implied conditions are read at face value.

Instead ofWrite
"Is this customer unhappy?""Does the customer say they are dissatisfied with the product or the service?"
"Is this a sales lead?""Does the sender ask about pricing, a demo, a quote or a start date for our product?"
  • Put the boundary cases in the criteria: option descriptions, score levels and the probability's yes / no.
  • When you catch yourself explaining a wrong answer ("but I meant…"), that explanation is the missing half of the instruction.
  • When interpretation is unavoidable, split it into two literal questions and combine them with an and condition or a composite.

2. Keep math, counting and dates in code

Jev is not a calculator. Compute in your application and send the result in the state — a number or, better, a named bucket.

{ "invoice": { "days_overdue": 42, "amount_band": "10k_to_100k" }, "message": "…" }
  • Counting: don't ask "how many items are fruits?". Ask one yes/no question per item — or one decision per item — and add up the answers in code.
  • Dates: don't ask which of two dates comes first. Extract the parts you need with choice questions (month, year, an explicit not_stated option), then compare and compute in code.
  • Names beat codes: "dark red" works better than #8B0000; convert in code first.
  • Weighted levels are for thresholds: scores tells you whether an answer is above or below a level, not the exact number between two levels.

3. Reduce indirection

Double negatives, "a property of a property" and multi-hop reasoning cost accuracy. Ask directly, and point at the part of the state the question is about by name, between backticks — dot and index paths work:

{
  "key": "refund_requested",
  "type": "probability",
  "instructions": "Does `ticket.messages[0].text` explicitly ask for a refund or a credit?"
}

4. Send only what the question needs

Unrelated detail in the state is a distractor: accuracy falls as the state grows with content the decision doesn't need, and every extra token counts toward the limits — about 32,000 tokens for the state plus the longest question and 64,000 for the state plus all questions.

  • Filter in your code first. Fields the schema doesn't declare are still forwarded to the engine, so don't send whole records "just in case".
  • Pass the relevant slice: the last messages of a conversation, the retrieved passages that passed a relevance check, the fields of the row the question is about.
  • Use a list state for transcripts and batches, and an object with descriptive keys for everything else.

5. Check the option order

Jev can lean toward the option that comes first in a choice question. Make the options mutually exclusive, then reorder them in the Playground and check that the answer stays the same. The reserved other option is always last, whatever the order of the others.

6. Don't ask it to generate

Jev returns choices, levels and probabilities — never free text, and neither does Dcision. For extraction, find the candidates in code (a regular expression, a parser or an LLM) and let a choice pick the right one. Option values are snake_case keys, so put each candidate in its description:

{
  "key": "customer_name",
  "type": "choice",
  "instructions": "Which candidate is the organization the invoice was issued to, as written in `source_text`?",
  "options": [
    { "value": "candidate_1", "description": "Beaver Logistics" },
    { "value": "candidate_2", "description": "Beaver Dam Logistics" },
    { "value": "candidate_3", "description": "Dam Logistics" }
  ]
}

7. Align instructions and criteria

Criteria extend the instruction; they shouldn't contradict it. A probability whose yes criterion describes the no case performs worse. Write what an average reader would understand at first glance.

8. Treat the state as data

The state is content, not instructions — and content can be written to steer the answer: an injected instruction, a misleading framing, text that argues for its own classification. Be explicit in the criteria about what counts, and run adversarial samples through the Playground before you deploy.

9. Use structure when it helps

Instructions, option descriptions, score levels and probability criteria accept JSON as well as text. Structure helps when a question has several parts or needs supporting data — the keys label each part. This choice describes what each option covers and doesn't cover:

{
  "key": "department",
  "type": "choice",
  "instructions": {
    "question": "Which team should handle this message?",
    "focus": "Classify the customer's primary request, not every topic mentioned."
  },
  "options": [
    {
      "value": "billing",
      "description": { "what": "Charges, invoices, refunds, or subscriptions", "not_for": "Order tracking or account access", "examples": ["I was charged twice", "Where is my refund?"] }
    },
    {
      "value": "orders",
      "description": { "what": "Order status, delivery, cancellation, or returns", "not_for": "Charges or account access", "examples": ["Where is my package?", "Cancel my order"] }
    },
    {
      "value": "account",
      "description": { "what": "Login, password, profile, or security", "not_for": "Charges or delivery", "examples": ["I can't log in", "Change my email"] }
    }
  ]
}

A score level can be a rubric too, as long as it has a label — the label is what result returns:

{ "label": "critical", "rubric": "outage, data loss, security issue or money at risk" }

See Structured instructions and criteria for the rules and limits.

10. Test in the language you serve

English is Jev's primary training language and where accuracy is best today. Other languages work, but test them on your own content and watch confidence when routing.

Checklist

  • One judgment per question, phrased as a literal question about the state.
  • Boundary cases and exclusions in the option descriptions, levels or criteria.
  • Numbers, counts and dates computed in code and sent as fields or buckets.
  • State fields named in the instructions with backticks when it removes ambiguity.
  • Only the state the questions need.
  • Options mutually exclusive, and the answer stable when you reorder them.
  • Adversarial and edge-case samples run in the Playground.
  • A minConfidence or a confidence rule where a wrong answer is costly — see Confidence-gated routing.

On this page