TypeSafe AI is building models for decisions inside software rather than conversations with people. Its first public model, Jev, accepts unstructured program state and a set of tightly defined questions, then returns typed answers with probabilities and confidence scores. The aim is to give developers an intelligence component that behaves more like a dependable function call than a free-form chatbot.
What Jev does
Jev is the first release in TypeSafe’s proposed “System One Model” category. Instead of generating an open-ended string token by token, it makes structured decisions in parallel. Developers define the possible output shape in advance and can ask several questions about the same state in one API request.
The current API supports three main decision primitives: yes-or-no judgments, selection from a fixed set of choices, and scores on a defined scale. Each result includes a probability distribution or confidence measure, allowing the surrounding application to set its own threshold for acting automatically, requesting review or taking a safer fallback path.
- Typed outputs that conform to the question schema
- Calibrated probabilities and confidence with every decision
- Multiple independent questions evaluated in one request
- Yes/no, multiple-choice and scored-rating primitives
- A REST API with an OpenAPI specification and model-discovery endpoint
- Designed for classification, routing, extraction, scoring and workflow branching
- Suitable for verification and guardrail steps around other AI systems
Where it fits
TypeSafe is aimed at developers building operational software: expense-policy checks, support routing, document triage, risk flags, quality scoring and other places where a program needs a narrow judgment before choosing its next action. Its low-latency design may also suit real-time interfaces or large data-processing pipelines where repeatedly calling a reasoning model would be too slow or expensive.
Jev is deliberately less general than a large language model. It does not write prose, hold a conversation or generate code. Applications still need ordinary code to define the workflow, enforce business rules and decide what to do with the returned probabilities.
Performance and reliability
TypeSafe reports end-to-end response times of roughly 70–500 milliseconds and, on its published workflow evaluations, gains as high as 193.6 times faster and 444.6 times cheaper than compared LLM workflows. Those headline comparisons come from TypeSafe’s own evaluation harness and the company says they are likely near the high end of real-world gains. Teams should validate Jev with representative private examples before using it for consequential automation.
The company describes Jev as having zero hallucinations because its output cannot violate the declared type. That is a guarantee about output structure, not a guarantee that every decision is correct. Jev can still make a wrong judgment; its practical advantage is that it exposes uncertainty so software can choose when to escalate.
Pricing and access
Jev is priced at $42 per billion input tokens, equivalent to $0.042 per million input tokens. Output is currently free because TypeSafe considers it too inexpensive to meter. The model is in early access, with developers admitted from a waitlist, so availability and commercial terms may evolve as the service moves toward broader production use.
Data handling
TypeSafe states that it does not train or fine-tune models on customer prompts or other input, and that it does not sell personal data or share it for cross-context behavioural advertising. It also publishes a data-processing addendum and a public trust centre. Enterprise teams should still review retention, hosting location and contractual controls for their particular workload.
Our take
TypeSafe’s strongest idea is not simply that Jev is fast; it is that many automation tasks need a constrained decision rather than another paragraph of generated text. The typed API and calibrated confidence make that idea concrete. Jev is promising for high-volume classification and branching, but it is an early product with a specialised interface and mostly first-party performance evidence. It is best evaluated as a component inside a carefully designed workflow, not as a drop-in replacement for every LLM call.