Jev is TypeSafe's structured decision model. It returns bounded judgments that application code can use when generated prose is unnecessary. For example, Jev can act as an LLM judge that classifies or scores support responses against defined criteria, making it useful for repeated evaluation where its measured quality, latency, and cost fit the application.
Braintrust supports Jev as an evaluation model, so you can point an existing scorer at it and apply the same criteria in experiments and online scoring. Start free with Braintrust.
What is Jev?
Jev is an AI model from TypeSafe built to make fast, structured decisions inside software. It understands natural-language input and answers questions whose possible outcomes you define in advance. An application might use Jev to identify the intent of a support request, rate the quality of an AI-generated response, or decide whether a document is relevant to a search.
Think of Jev as a fast decision step or classifier. It bridges a gap between traditional ML, where you create a specific classifier model for one task, and LLMs, which return free-form answers. Jev is effectively a general-purpose classifier with a world model and tuned weights.
When you're looking for answers that depend on understanding language but fit into neat classifications, such as whether a customer wants a refund or whether a response answers their question, Jev is a good fit for making those determinations quickly, cheaply, and without straying outside the allowed answers.
TypeSafe calls this category System One, referring to fast, focused judgments.
How Jev answers prompts
To ask Jev a question, you provide two things:
- The information it should read
- The kind of answer you want
The information it reads is known as its state. This could be a customer's message, a draft reply, or account details needed to check that reply.
You then write your question and choose an answer format. For example, you could give Jev a support ticket and ask which team should handle it. You supply the choices — billing, technical support, or other — and Jev returns its selection along with a probability for each option.
Jev supports three question types:
| Question type | What you provide | What Jev returns | Example |
|---|---|---|---|
| Choice | A question and a list of possible answers, up to 255 options. | The selected answer, a probability for each option, and confidence. | Which team should handle this ticket: billing, technical support, or other? |
| Score | A question and an ordered list of 2 to 10 rating levels described in words. | A score, probabilities for the levels, and confidence. | How complete is this reply: missing the answer, answering part of the request, or answering it fully? |
| Noul | A yes/no question, with any details needed to define what counts as yes. | A number from 0 to 1 showing how likely Jev thinks yes is. | Does this reply promise a refund that the account policy does not allow? |
For a Choice question, an illustrative result might give billing an 80% probability, technical support 15%, and other 5%. Jev would select billing.
For a Score question, the levels are numbered starting at zero, so three levels give you a scale from 0 to 2. The returned value is the probability-weighted mean of the level indices, which means Jev can return a value between levels, such as 1.4. That reflects how it distributed probability across the levels. It does not mean the reply is "70% complete."
For a Noul question, which TypeSafe names as a portmanteau of "boolean," 0.9 means Jev estimates a 90% probability of yes. A value near 0.5 means it is unsure. Noul expresses its uncertainty through that probability rather than a separate confidence value.
Choice and Score also return confidence, which summarizes how strongly the probabilities favor one answer. Your code can use these results to send a ticket to a team or flag a reply for review. Check them against real examples before deciding when to act automatically.
Jev accepts text, including text organized in JSON objects or lists. It cannot read images, audio, or video, or generate replies, code, or explanations.
How is Jev different from other LLMs?
Jev focuses on bounded semantic decisions rather than free-form generation. Generative LLMs can also classify inputs and return valid JSON, but Jev uses a decision-focused interface and training objective built around typed questions and structured answers. Application code receives probabilities and predefined labels instead of prose or reasoning.
Jev evaluates independent questions against shared state in parallel. For example, one request can check whether a support reply answers the customer and whether it makes an unsupported claim. One answer cannot feed into another question in the same request, so application code must handle dependencies, calculations, control flow, and actions. TypeSafe says additional questions barely change response time, though every question adds input tokens.
TypeSafe calls Jev's training approach Reinforcement Learning for Calibrated Decisions (RLCD). Calibration describes how predicted probabilities behave across groups of examples. For Choice and Score, confidence measures how concentrated the probability distribution is, not the chance that the judgment is correct. Jev can return a valid schema with the wrong allowed answer, and constrained outputs do not protect it from prompt injection: TypeSafe's model jaggedness notes for Jev 1.13 state that the model does not treat state as hostile by default, so injected instructions or deliberately misleading framing can move the answer.
As of September 2026, direct Jev pricing was $0.042 per million input tokens, with free output. TypeSafe reported response times of 70 to 500 milliseconds, but actual latency depends on the request and workload. These economics can suit repeated judging or routing when Jev's measured decision quality meets your requirements.
What tasks is Jev best at?
Jev fits repeated, bounded evaluation tasks such as LLM-as-a-judge. Braintrust supports Jev judge scorers for experiments and online scoring. For example, you could give Jev a customer request and verified account facts, along with a draft support reply. Separate questions could check whether the reply makes unsupported claims and whether it answers the request. A Score question could assess the completeness of the proposed next steps. Low latency and token cost make these narrow checks practical across large sets of outputs or traces.
You still need to evaluate Jev as a judge. Build a test set with human labels, ambiguous examples, and realistic failures. Measure false approvals and false rejections separately. Track how often your policy escalates a case for review. Compare judgment quality with your current judge, then compare latency and total workflow cost on the same examples. Observed errors should determine thresholds rather than an arbitrary confidence cutoff.
Jev also fits routing and classification when the possible destinations are known. It can select a ticket category or handler, while application code performs the lookup and dispatches the request. TypeSafe documents this division in its intent-routing pattern.
RAG and agent evaluation offer similar bounded questions. Jev can classify a retrieved passage as relevant or conflicting before another model writes an answer. It can also inspect an agent trace for a defined failure, while your application collects the trace and applies the review policy. TypeSafe provides examples for RAG passage classification and agent-trace review.
Use a generative model when your application must write or summarize text. Coding and extended reasoning also fall outside Jev's intended role. Exact arithmetic and date comparisons belong in ordinary code, since TypeSafe documents weak arithmetic, counting, and date ordering among Jev 1.13's failure modes, alongside degradation from irrelevant context and no default hostility toward adversarial input. Typed outputs prevent malformed responses, but Jev can still select the wrong allowed answer.
How can people use Jev?
There are three ways to get started:
- Use Jev in Braintrust. Use Jev to check and score AI responses. You can connect a TypeSafe API key or request built-in Jev access through Braintrust.
- Join TypeSafe's waitlist. Sign up on TypeSafe's website for direct access to Jev.
- Use OpenRouter. Access Jev 1.13 through OpenRouter to use it in your application.
Using Jev in Braintrust
In Braintrust, you can use Jev as a scorer — a check that labels or rates an AI response. Here's how to set it up:
- Add TypeSafe under AI providers using your TypeSafe API key. To use Jev without a separate TypeSafe key, request built-in access from Braintrust, then select Jev from Braintrust's built-in models.
- In the scorer, select Jev as the model.
- Write what Jev should check and include the information it needs, such as the customer's request, verified facts, and the AI's reply.
- Under Classifications or Scores, define the possible answers. For example, use "Answers the request," "Missing key details," and "Makes unsupported claims."
- Run the scorer on test examples and inspect its answers, confidence, and probabilities. Once you've checked its quality, you can also use it in online scoring to apply the same criteria to production traces.
Braintrust saves the selected label along with the model that returned it, token usage, and duration, and records confidence and per-choice probabilities in metadata so you can review the results. Braintrust's Python and JavaScript libraries can also record TypeSafe calls made directly from your application, capturing each one as a typesafe.systemOne span with its inputs, questions, answers, and token usage.
See Eval agent responses with Jev for screenshots and examples.
Frequently asked questions
What kinds of questions can Jev answer?
Jev answers multiple-choice questions, scores inputs against a defined rubric, and estimates the probability that a yes or no statement is true. Those correspond to its Choice, Score, and Noul primitives. Jev does not generate free-form responses.
How much does Jev cost?
As of September 2026, Jev costs $0.042 per million input tokens, with no charge for output tokens. Separate platform and workflow costs may still apply.