How Do Jev, LLM Routing, and Guardrails Work Together?

Jev decides closed-choice probabilities; code routes to LLM/human and enforces guardrails, not LLM. TypeSafe official intent-routing/RAG/guardrail examples; PandaNpc calibrates 19 synthetic turns with mock LLM provider.

PandaNpcFirst published on
How Do Jev, LLM Routing, and Guardrails Work Together?

Disclosure of Interests and Evidence: PandaNpc is developing an Agent decision layer that uses Jev. Below, we cite TypeSafe official documentation, official cookbooks, and real Jev calls and synthetic-scenario calibration from our repository. Our calibration uses a scripted fake LLM provider and cannot represent real user traffic or the performance of a complete production pipeline.

LLM routing can work like this: first let Jev determine what category a request belongs to and how risky it is, then let code decide whether to hand it to an ordinary function, a specialized LLM, or human review. Jev can also sit between retrieval and generation to filter evidence, or after LLM output to check results. It returns closed options, scores, and probabilities; open-ended answers, code generation, and long reasoning are still handled by the LLM. TypeSafe's explanation of coding agents clearly states that Jev cannot directly replace the chat model behind Claude Code or Codex.

This article uses a customer service request, a RAG question-answering pipeline, and our own Agent calibration records to explain where the two types of models actually hand off to each other, and why low-confidence results must have a definite destination.

What Can the TypeSafe Jev Model Judge, and What Does the LLM Continue to Handle?

As of September 23, 2026, the stable model listed on the TypeSafe model page is jev-1.13.0. The API accepts a state and a set of questions, and returns the corresponding structured answers via POST /v1/systemone. jev-latest pointed to 1.13.0 that day, but aliases change with versions; systems whose thresholds have been calibrated should pin the version and record the actual model ID in the response.

Question type What it is good for asking What it returns What code does with it
Choice "Is this request a refund, an order lookup, or a complaint?" One fixed option, probabilities for each option, confidence Decide target handler; escalate low confidence
Score "What level of severity is this complaint at?" Level score, probabilities for each level, confidence Compare with business thresholds
Noul "Is the user explicitly asking for a refund?" Probability of "yes", 0–1 Set allow, reject, and pending-review ranges based on probability

Noul has no separate confidence field; you cannot directly write a Noul probability as "model confidence." Score also should not be used to calculate exact amounts. Amounts, date comparisons, quotas, and permission checks should remain in deterministic programs; the official documentation lists these boundaries for Jev 1.13.

Original flow diagram from input to Jev's three question types, code thresholds, and LLM or human review
Original diagram: a request goes through Jev to produce closed answers, and code decides the next step according to this system's thresholds; the arrows only indicate one possible architecture, not a product interface or measured results.

The Minimal Shape of a Single Call

The request shape below is consistent with the official API reference; the example questions are an illustrative configuration constructed for this article, and we did not run it online for this article:

json
{
  "model": "jev-1.13.0",
  "state": {
    "message": "My order was charged twice. Please help me get a refund.",
    "account_note": "Customer asks about an order charge"
  },
  "questions": {
    "intent": {
      "type": "choice",
      "instructions": "What does state.message primarily request?",
      "criteria": {
        "refund": "Money returned for a charge",
        "information": "An explanation only",
        "other": "Neither option fits"
      }
    },
    "asks_refund": {
      "type": "noul",
      "instructions": "Does state.message explicitly ask for money back?"
    }
  }
}

An actual system should also first use code to check whether the charge records belong to the same order and whether a refund is allowed. The example above is only for interpreting user intent; a user asking for a refund does not mean refund eligibility has been proven, let alone constitute authorization to directly execute a refund.

Using Jev for LLM Routing: Three Handoff Paths

TypeSafe's official intent routing example first hands a customer service request to Jev to determine intent and complexity, then code routes it: order status lookups go to a database function; product questions and returns/exchanges go to specialized LLMs loaded with different materials; complex complaints or low-confidence results enter a human queue. This is the most understandable form of Jev-LLM collaboration: the former gives structured judgments, while the latter appears only when an explanation or conversation needs to be generated.

When implementing, you can design in the following order instead of letting the model freely decide all actions:

  1. Define paths first: List clearly which requests ordinary functions, each specialized LLM, and human review can handle, and leave other or a similar fallback option for Choice.
  2. Put facts into state: Make the user's exact words, account status, and order records into separate fields; do not treat unattributed web page text as system instructions.
  3. Ask narrow questions one at a time: Use Choice for intent, Score for risk or urgency, and Noul for a single fact that needs confirmation. The official guidance suggests that multiple independent questions sharing the same state can be evaluated in parallel in the same request.
  4. Let code perform final routing: First check permissions and hard rules, then look at Jev's probabilities and thresholds calibrated for this business; requests with low confidence or missing evidence go to human review or a follow-up question.
  5. Record results and review them: Save the model version, question version, probabilities, final destination, and human correction results so you can judge whether the thresholds are appropriate.
Chart of evaluation metrics and per-workflow cost for model results relative to reference probabilities across four TypeSafe-built workflows
TypeSafe official chart: metrics and cost for four self-built workflows; accuracy references the average predicted probability from GPT-6 Astra and Claude Fable 5.1, not human ground-truth accuracy, and not a PandaNpc measurement.

Image source: TypeSafe AI, "Introducing System One Models & Jev", 2026-09-15. The four workflows were built by TypeSafe, and metrics are aggregated with equal weight per workflow; for the evaluation method, see TypeSafe workflow evals.

This official chart helps explain why it emphasizes "packing multiple narrow judgments into a program workflow." The vertical axis uses the vendor's "accuracy" label, but its reference answers come from a consensus of predicted probabilities from two large models, not a single manually verified correct answer; cost and metrics also depend on these four workflows and the vendor's evaluation method, and cannot be converted into "how much can be saved in any scenario."

Retrieval-Augmented Generation: Jev Filters Evidence Before the LLM Answers

TypeSafe's RAG passage cookbook provides a more concrete multi-model example: OpenAI embedding first retrieves passages, and Jev asks four Noul questions for each "question + passage"—whether it is relevant, whether it contains evidence usable for answering, whether it contradicts a premise in the question, and whether it attempts to instruct the answering model. Code processes the four probabilities in order to decide whether to place the passage in the evidence section, the conflicting evidence section, or discard it; finally Claude Sonnet 5 writes the answer.

This step solves a common problem: passages with high vector similarity are not necessarily usable. They may merely use similar words, or they may be a forum post carrying a prompt injection like "ignore the preceding text." The cookbook example places the injection check first in the routing rules, while reminding readers that the thresholds are a starting point selected for that corpus, not default values for all RAG applications. Its demonstration numbers come from jev-1.12 on 2026-08-27 and should not be treated as new evaluation results for the current jev-1.13.0.

Original RAG flow diagram: retrieved passages pass through Jev checks for relevance, evidence, conflict, and injection before entering the LLM answer
Original diagram: four narrow judgments jointly decide whether a passage stays or goes; actual thresholds need to be validated with your own corpus.

After generation, you can also add a layer of verification. TypeSafe's citation check cookbook first uses a program to find the cited source text, then uses Jev to judge whether that passage supports, contradicts, or does not mention the generated claim. It can surface citations worth reviewing; the model judgment itself can still be wrong, and "passed the check" should not be written as a factual guarantee.

Our Agent Calibration: Where Do Low-Confidence Escalations Get Stuck?

In the PandaNpc repository, the Jev client, question bank, and orchestrator use Jev for Agent intent recognition, candidate edit scoring, completion condition checking, and submission decisions. The client also does limited retries for timeouts, 429s, and 5xxs, and sets limits on request budgets and expired results; execution permissions are held by the orchestrator and controlled tool layer, and are not directly granted by a single Jev judgment.

On 2026-09-22, using jev-1.13.0, we ran one shadow and one enforce pass for each of 19 synthetic turns, 38 runs in total, and recorded 165 real Jev decisions. This internal calibration report and the saved real-machine responses used a scripted fake LLM provider, so these data only show decision performance in a controlled scenario. They cannot prove overall success rate, savings ratio, or end-to-end latency under real user requests.

The most valuable finding was not average speed, but a blockage caused by a threshold that "looked safe": among the 19 enforce turns, 13 escalated at Q2, "is there enough information to start editing?", because the Noul probability fell into the originally planned 0.15–0.85 uncertain interval; the LLM Worker never had a chance to execute subsequent steps. The calibration records show that among 34 Q2 decisions labeled as having sufficient information, many probabilities were in the middle range. The report recommends splitting the complex Q2 into more atomic judgments, or adjusting the escalation rules; these are recommendations, not thresholds that have already been deployed.

Our question bank classifies the entry point as answer_only, inspect, modify, or out_of_scope; on the write path, candidate edits proposed by the Worker are first ranked by Score, and the final content and change summary then go through acceptance and submission decisions. These are only decision points: whether objects can actually be read or written is still governed by the controlled executor granting permissions by phase. Jev has no authority to relax the tool allowlist itself, nor can it bypass the pre-submission consistency check.

The calibration data reveals another trade-off. In shadow mode, Jev gives answers and full distributions but does not change the Worker's original execution path; in enforce mode, the answers affect whether to continue, escalate, or discard. Treating shadow accuracy directly as enforce completion rate misreads the system: Q2 escalation cuts off tasks early, so subsequent candidate scoring, acceptance, and submission questions never get a chance to appear. Therefore, this report reads per-question distributions, escalation direction, and final state separately.

There is a concrete contrast among candidate edits: in the same turn, a candidate that precisely edits the target scored 2.94, while a candidate that overwrites the entire file scored 0.38; the higher-scoring candidate was selected. This example only shows that, in that synthetic situation, the scoring question distinguished the two options. Conversely, a candidate with truncated evidence scored 2.27; you should not ignore its low confidence and truncation flag just because the number looks "not bad." Our code marks incomplete evidence separately to prevent the model from making a deterministic write decision based only on the retained prefix.

We also split "which candidate to choose" and "allow it to write" into two different steps. After a candidate is scored by Jev, the controlled executor exposes write tools only during the ACT/modify phase; the issued one-time ticket is bound to the tool call ID, current revision number, target object hash, and parameter digest. Even if the candidate text induces the model to "ignore restrictions," it cannot obtain tool permissions that bypass these checks. This is our experience from code-level integration: probability judgments decide which path is worth taking, while side-effect permissions are determined by reviewable program conditions.

Failure paths must also be designed. The client retries only in a limited way on timeouts, network errors, 429s, or 5xxs; answers after cancellation or past the turn deadline are discarded directly. If Jev is unavailable in enforce mode, the system cannot continue without downgrade authorization; when authorized to use llm_only, the executor is locked to read-only. If a branch has already been modified before Jev is lost, the orchestrator marks the entire turn as failed, rather than letting a subsequent LLM fill in the write without the decision layer. These paths have a cost in user experience, but they keep "the model is temporarily unavailable" from quietly becoming "write permissions as usual."

The calibration report also distinguishes "low-confidence escalation" from "refusal to execute." For example, a correct discard whose confidence does not reach the uniform 0.85 threshold is recorded as needing user input; this does not equal an erroneous allowance. On this basis, the report recommends separating the thresholds for submission and discard, though these are still recommendations for now. When writing workflows, you must distinguish three outcomes: erroneous allowance, erroneous rejection, and pending review; otherwise the same dataset will lead to wrong threshold conclusions.

This case tells us that Jev-LLM collaboration cannot be drawn only as "Jev judges first, LLM works afterward." For every judgment, you must ask: How wide is the uncertainty interval? Will it prevent downstream processors from ever receiving the task? If the input evidence is truncated, can it clearly escalate instead of guessing? In our implementation, the state builder records evidence_truncated and makes the caller treat the missing-evidence path as uncertain; deterministic computations such as counting and sorting are done in code first, not left to Jev to guess. The official Jev 1.13 known limitations also recommend keeping counting and arithmetic in code.

Where Should LLM Guardrails Go?

TypeSafe's LLM guardrails cookbook places Jev on both the input and output sides of the LLM. It uses a set of Noul questions to identify different risks, uses Score to measure severity, and then code decides according to policy whether to allow, manually review, block, or hand off to support. Output must also be checked, because ordinary input can still produce inappropriate generated results.

The boundaries of this kind of guardrail are equally clear: Jev can check content according to prewritten questions, but it is not a universal security proof. The official limitations document explicitly mentions that malicious content may influence judgments, and requires writing clear criteria and testing boundaries. In our synthetic samples, we ran 16 injection probes targeting candidate parameters and recorded 0 ranking flips; the sample is too small to conclude that "prompt injection resistance has been solved." What actually determines what a tool can do is still the allowlist, phase gates, and pre-submission checks in code.

When Is It Appropriate to Use It, and When Should You Not?

Jev is suitable when: the candidate set is known, the question can be split into several short judgments, and the software needs probabilities to decide between automatic handling and escalation to a human. Examples include customer service routing, RAG passage filtering, Agent candidate action scoring, and citation checking in generated results. If the task requires writing a reply, editing a piece of code, or explaining a complex reasoning process, the LLM takes over. If the task is to calculate money precisely, compare dates, or check access control, the program should compute directly. The model page also states that Jev accepts only text, and English is currently the training language with the best performance; Chinese-language scenarios need to be evaluated with your own data and cannot copy the thresholds from English cookbooks.

If you want to observe how a real Agent handles permissions and tool calls, start with PandaNpc Agent; for the boundaries and use cases of coding agents, you can also see Claude Code vs Codex.

Readers can start with a very small validation set: prepare four categories of samples—"clearly can be handled automatically," "clearly should be rejected," "semantically ambiguous," and "contains malicious instructions"; first determine human labels, then record Jev's per-question probabilities and routing results. The success criterion is not that every item passes automatically, but that the error rate of the automatic handling path and the amount of human escalation both fall within a range you can accept. If many ambiguous samples get stuck on the same question, first check whether the question mixes multiple judgments, whether the state is too long, or whether the thresholds were calibrated on local data.

FAQ

Can the TypeSafe Jev Model Replace Claude Code, Codex, or a Chat Model? No. TypeSafe positions it as a structured decision model inside software; chat, writing, and code generation still require an LLM.

Does a fixed return type mean it will not make mistakes? No. Fixed types reduce parsing and out-of-bound output problems, but classification, scoring, and factual judgments can still be wrong. Low-confidence and high-risk paths should retain human review.

How many questions can one request ask? You can put multiple independent Choice, Score, and Noul questions that share the same state in one request. Each question is evaluated separately, and complex judgments should still be split apart and then combined by code.

Can Chinese be used? The official documentation says it supports natural languages including Chinese, Japanese, and Korean characters, but English accuracy is currently the best. Chinese-language workloads need separate validation and calibration.