Blog

System 1 AI Models: Jev, Laya, and the Rise of Decision AI

By Gourav Bais . Tech Lead

Having trouble scaling AI across your business?

Get in Touch

Having trouble scaling AI across your business?

Get in Touch

A technical guide to TypeSafe AI’s decision model, its open-source rival Laya, and where this new model category fits in production

Jev is TypeSafe AI’s first “System One” model – a schema-constrained AI that returns typed, calibrated decisions instead of text. Rather than generating a written response one token at a time the way a chat model does, it takes a piece of state (an email, a support ticket, a JSON object) plus a set of typed questions you define, and returns a value from that schema, a choice, a score, or a probability, alongside a calibrated confidence score, computed in a single non-autoregressive pass.

TL;DR  Jev and its open-source rival Laya belong to a new, fast-growing category of AI model built for structured decisions, not conversation. Jev launched September 15, 2026, from TypeSafe AI, built by ex-OpenAI researcher Diogo Almeida; it reports 70–500ms latency at $0.042 per million input tokens, output free.An open-weights alternative, Laya, followed within about a week from Convai Innovations, a 421-million-parameter ModernBERT-based model, free to self-host under Apache 2.0.Neither model generates text. Both answer typed questions (a choice, a score, or a yes/no probability) against a state you provide, in one parallel pass.Real production use cases so far: cost-aware LLM routing, document and ticket classification, conversation-memory triage, and pre-action safety checks for autonomous agents.The catch: strong zero-shot accuracy isn’t guaranteed for either model. Production-grade results generally require fine-tuning and confidence recalibration on your own data.

What Is a “System 1” AI Model?

Every AI model most people talk about in 2026 GPT, Claude, Gemini is what TypeSafe AI calls a “System 2” model. Give it a prompt, and it generates a response one token at a time, threading together a chain of reasoning into readable prose. That’s genuinely useful for writing, chatting, and multi-step problem solving but it’s a slow, expensive way to answer a question that only has a few possible answers, like “is this email spam?” or “which of our three support queues should this ticket go to?”

A System 1 model skips the “write an answer” step entirely. A System One model is a frontier AI model that returns typed probabilistic decisions instead of text, as one glossary puts it: you give it a block of state and a set of typed questions, it evaluates all of them in parallel, and it returns structured answers with calibrated confidence scores, nothing to parse, and no string generation at any point.

The category name is a deliberate reference to Daniel Kahneman’s Thinking, Fast and Slow, which splits cognition into System 1 (fast, intuitive judgment) and System 2 (slow, deliberate reasoning). A System 1 model is built for the kind of judgment call a piece of software makes thousands of times a day — not the kind of task where you actually want to read an explanation.

System 2 (generative chat models) vs. System 1 (decision models like Jev and Laya)

Meet Jev: TypeSafe AI’s System One Model

TypeSafe AI came out of stealth on September 15, 2026, with $40 million in seed funding led by DCVC, and a single product: Jev. Founder Diogo Almeida is an OpenAI veteran and co-inventor of RLHF (Reinforcement Learning from Human Feedback) the technique behind InstructGPT and, later, ChatGPT and he built Jev alongside co-founders Erik Gafni and Sasha Sheng.

TypeSafe describes Jev’s architecture as a parallel sampler: rather than generating text token by token, it processes an unstructured state plus a set of typed questions, and computes every required answer in a single parallel pass, with each question in a call evaluated independently against one shared read of the state. The name is a double reference “Jev” nods to 19th-century economist William Stanley Jevons, whose paradox holds that making a resource cheaper increases total demand for it rather than reducing it. TypeSafe is betting that sufficiently cheap, fast decisions get wired into far more places than a slow, expensive LLM call ever could.

The model answers exactly three kinds of typed question, which can be combined and run in parallel within a single API call:

  • Choice – pick one option from up to 255 alternatives (larger option sets use a two-pass score-then-choose approach)
  • Score – rate the state against an ordered rubric
  • Noul – return a calibrated probability that a yes/no statement is true

Every answer comes back with a per-option probability distribution and an aggregate confidence score between 0 and 1, derived from the shape of that distribution – so a “billing” classification that wins with 84% probability but leaves a meaningful share on a competing option produces a lower confidence score than a 99%-to-1% split, even though both technically “chose billing.”

Speed, Cost, and the “0% Hallucination” Claim

TypeSafe’s headline numbers: 70–500 milliseconds end-to-end latency, versus a reported 3 to 329 seconds for the frontier LLMs it benchmarks against, at $0.042 per million input tokens with output free (there’s nothing to charge for, since no output tokens are generated). On the company’s own four-workflow production benchmark, Jev scores around 67.8% agreement – close to GPT-5.6 Terra’s 67.9% on the same tasks — which TypeSafe prices at roughly $0.39 versus $3.31 per thousand workflows run through a comparable LLM.

The more provocative claim is “0% hallucination.” That’s true in a narrow, specific sense: because the answer space is defined by the schema passed in at request time, Jev cannot return malformed output or a value outside the options offered, there’s no free-text generation step where a model can invent a plausible-sounding wrong answer. What it can still do is confidently choose the wrong valid option. TypeSafe trains for this with what it calls RLCD (Reinforcement Learning for Calibrated Decisions), in place of the RLHF approach used to tune conversational models optimizing for confidence scores that are actually trustworthy against strictly proper scoring rules, not just for getting the top answer right.

Under the Hood: How a Decision Model Actually Works

TypeSafe hasn’t published Jev’s architecture, but the general shape of a System 1 decision model is well understood and it’s exactly what open alternatives like Laya expose publicly:

  • Input state. The process starts with whatever you want a decision about a sentence, an email, a support ticket, a conversation transcript, or a JSON object.
  • Tokenization. Like any transformer-based model, the input is split into tokens and mapped to numerical IDs, since neural networks operate on numbers, not words.
  • Bidirectional encoder. This is the backbone. Unlike the decoder-only architecture behind chat models which only look backward at tokens already generated a bidirectional encoder from the BERT family reads the entire input at once, so every token’s representation is informed by every other token around it, not just the ones before it.
  • Decision head. A small neural network sits on top of the encoder, trained specifically to convert that contextual representation into the requested output type a class, a score, or a probability.
  • Typed, calibrated output. The final answer returns as a value plus a confidence score, never as generated text.

This isn’t conceptually new teams have paired encoder embeddings with a classification head (logistic regression, a small neural net, even a random forest) for years, going back to the ResNet/VGGNet era of computer vision, where a pretrained network’s final layer was stripped off and its remaining embeddings were fed into a separate classifier. What’s changed is the backbone: instead of a general-purpose image or language embedding model, today’s decision models pair the same technique with an encoder that already understands natural language at something closer to frontier-model quality, and wrap the whole pipeline in an API that returns a calibrated confidence score by default.

The general System 1 decision-model pipeline, shared by Jev and Laya

Meet Laya: The Open-Source Alternative

Roughly a week after Jev’s launch, Convai Innovations released Laya an open-weights, Apache 2.0-licensed System 1 decision model available on Hugging Face and via pip install laya, positioned explicitly as the open answer to TypeSafe’s closed API.

Laya’s architecture is public, which makes it a useful stand-in for understanding what Jev is probably doing internally. The main English checkpoint pairs a ModernBERT-large encoder (28 layers, 395 million parameters) with a roughly 25-million-parameter decision head two transformer layers, an option-marker scorer, and a small “answer-or-escalate” routing head for 421 million parameters in total. ModernBERT itself, released by Answer.AI and LightOn in December 2024, is a modernized BERT-family encoder trained on 2 trillion tokens of English and code, with a native 8,192-token context window, rotary positional embeddings, alternating global/local attention, and GeGLU activations architectural choices aimed squarely at making an encoder-only model fast and cheap to run at scale, which is exactly what a decision model needs from its backbone.

The clever part of Laya’s design is how it scores an open-ended set of options without retraining. Every candidate option is rendered as text with its own [MASK] token; the model reads the hidden state at that marker to produce a logit, then applies softmax across all the options belonging to that question. Because the answer space is defined by what’s rendered at request time not baked into the final layer’s weights adding a new category or a new question doesn’t require retraining the model at all.

Laya ships three checkpoints: an English base (ModernBERT-large, 421M parameters, 512-token context), a multilingual variant on mmBERT-base (322M parameters, 100+ languages), and a typed-decisions checkpoint fine-tuned on a public typed-decision benchmark. On that benchmark, the fine-tuned checkpoint reports 0.766 accuracy against Jev 1.13.0’s 0.727, at roughly 33ms latency versus Jev’s reported 236–276ms. Zero-shot performance, by contrast, is far weaker around 0.362, barely above a 0.318 random baseline and raw calibration is poor out of the box; refitting a single temperature value per question type reportedly brings mean calibration error down from 0.466 to 0.081. The honest summary: Laya is a fast, well-designed base to specialize on your own data, not a zero-shot decision engine you can drop in blind.

An Origin Story Worth Knowing

Laya’s release came with a pointed backstory. Convai Innovations is associated with researcher Nandakishor M, who published a paper titled SalesRLAgent in March 2025, describing a non-autoregressive, reinforcement-learning-trained model for predicting sales-conversion probability in real time reporting 96.7% accuracy and 85-millisecond inference versus roughly 3,450 milliseconds for GPT-4 on the same task. When Jev launched in September 2026 on a similar core idea typed decisions, calibrated confidence, no text generation without TypeSafe releasing its own weights or training data, Nandakishor published a response arguing he’d shipped essentially the same non-autoregressive, RL-trained approach well over a year earlier.

It’s worth being precise about what’s actually established here. The underlying idea training a bidirectional model with reinforcement learning to output calibrated decisions instead of text clearly predates Jev by well over a year in the published literature. Whether TypeSafe’s team was aware of that specific paper isn’t something either side has confirmed publicly, and TypeSafe hasn’t disclosed Jev’s architecture in enough detail to compare the two directly. What is clear is that the polished, Apache 2.0-licensed Laya package most people are installing today wasn’t around before Jev its own model card is dated within about a week of Jev’s launch, packaged and open-sourced squarely in response to the moment.

Reported latency and price ranges for Jev, Laya, and a typical frontier LLM

Jev vs. Laya vs. a Frontier LLM, at a Glance

 Jev (TypeSafe)Laya (Convai)Frontier LLM
License / accessClosed, hosted APIOpen weights, Apache 2.0Closed or open; hosted or self-hosted
OutputTyped choice / score / probabilityTyped choice / score / probabilityFree-form text
Typical latency70–500 ms33–40 ms (self-hosted, GPU)3–329 sec (reported range)
Price / 1M input tokens$0.042 (output free)$0.00 (self-hosted)~$0.20–$10+
ParametersUndisclosed421M (English checkpoint)Tens to hundreds of billions
Needs fine-tuning?Recommended for narrow domainsStrongly recommended (zero-shot is weak)Not required (general-purpose)
Best forLow-ops hosted routing / classificationSelf-hosted, data-sensitive, customizableWriting, chat, open-ended reasoning

Real-World Use Cases

1. Cost-Aware Model Routing

Sending every incoming prompt straight to a frontier reasoning model wastes money and adds latency on the requests that never needed it. A decision model sitting in front of the router can classify each prompt’s complexity — does this need Opus-class reasoning, or will a cheaper Haiku-class model do? — for a fraction of a cent, forwarding only the genuinely hard cases upstream. Because the classification call costs roughly $0.042 per million tokens against several dollars per million for a frontier model, the routing layer can pay for itself the first time it correctly downgrades a simple request.

A decision model used as a cost-aware router in front of multiple LLM tiers

2. Conversation and Context-Window Triage

As a conversation grows, an agent’s context window eventually fills up, forcing a decision about what to keep and what to summarize or drop. Instead of asking an LLM to make that judgment call in natural language — expensive, and itself consuming context — teams are using decision models to score each candidate chunk of conversation state against a simple typed question (“does this need to stay in active memory, yes or no?”), retaining only chunks above a confidence threshold. Because each chunk is scored independently within the same request, this scales to hundreds of candidate memories without a proportional latency cost.

3. Document and Ticket Classification

Classic multi-class or binary classification — sorting incoming support tickets into queues, tagging documents by category, flagging spam — is close to the textbook use case for this model family, and it’s also where the “it just works out of the box” claims run into the sharpest limits. Zero-shot accuracy on a domain-specific document set (legal filings, internal taxonomies, specialized product categories) tends to land well below headline benchmark numbers, since those benchmarks were built on different label distributions. Getting production-grade accuracy on your own categories typically means fine-tuning the decision head — and sometimes the encoder — on a labeled sample of your own documents. That’s not a large undertaking for a 421-million-parameter model, but it is a real step, not a zero-shot deployment.

4. Guardrails for Autonomous Coding Agents

One of the more concrete public examples of this pattern is gating what an autonomous coding agent is allowed to execute. Before running a shell command an agent has proposed, a decision model can be asked a typed question — run it, reject it, or ask a human? — alongside a boolean flag for whether the command looks irreversible (would it destroy data or leak a secret?). A command like rm -rf ./build comes back flagged for human confirmation because it’s irreversible, even though it might otherwise be an ordinary build-cleanup step — exactly the kind of narrow, repeated, structured judgment call this model category is built for.

The Broader “System 1” Model Ecosystem

Jev and Laya are the two names getting the most attention, but they’re not alone. In the weeks since launch, a cluster of similarly-shaped open models has appeared, most fine-tuned on top of existing open backbones rather than trained from scratch: decision-focused models built on Qwen3.5 for agent and tool-calling states, guardrail-focused releases from established model labs, general-purpose NLI-based classifiers, and lightweight community projects distributed through Ollama-style local runners built specifically for this model category. The common thread is the same interface Jev and Laya share — a state, a set of typed questions, and a typed answer with a confidence score — which suggests the industry is converging on a shared shape for this kind of model faster than it converged on, say, a shared plugin format for chat models.

Where System 1 Models Fall Short

  • Not for chat, summarization, or anything requiring a written explanation — that’s definitionally out of scope.
  • Zero-shot performance on unfamiliar domains is often close to a random or majority-class baseline; production-grade accuracy generally requires fine-tuning on your own labeled data.
  • Confidence scores are only as trustworthy as the calibration behind them — an unrecalibrated model can be significantly over- or under-confident on your specific data distribution.
  • Domain specialization matters. A general-purpose encoder wasn’t trained on medical or legal text at the depth a domain-specific model was; specialized fields typically need a domain-adapted encoder, not just a different decision head.
  • Sending sensitive or regulated data to any hosted third-party API — including Jev’s — carries the same data-governance considerations as any other cloud AI service; self-hostable models like Laya sidestep that specific concern by keeping inference on your own infrastructure.
  • Neither vendor’s benchmark numbers are independently audited yet. TypeSafe’s own comparisons use other frontier models’ outputs as a reference rather than independently verified ground truth, and Laya’s headline accuracy figure belongs to a checkpoint fine-tuned on that same benchmark’s training split.

Bringing System 1 Decision Models Into Production

Choosing between a hosted API like Jev, a self-hosted open model like Laya, or a custom-trained encoder-plus-classifier isn’t really a “which model is best” question — it’s an architecture decision that depends on data sensitivity, latency budget, the existing model stack, and how much labeled data is already on hand for fine-tuning. That evaluation work is exactly what AI engineering partners handle day to day.

This is the kind of work we do at 47Billion. Our AI/ML practice already covers the adjacent pieces a production decision-model rollout sits directly beside — transformer-based NLP for named-entity recognition and aspect-based sentiment analysis, intelligent document processing that combines OCR with classification, and enterprise AI agents for support and operations workflows. Slotting a Jev- or Laya-style decision layer in front of an existing LLM stack, fine-tuning it on your own labeled tickets or documents, and wiring its confidence thresholds into a routing or escalation workflow is a natural extension of that same practice, rather than a separate discipline. Whether the right move for your team is a hosted API, a fine-tuned open model behind your own firewall, or a hybrid of both depends on specifics no vendor benchmark can answer alone — which is exactly the kind of evaluation we run with clients. If you’re weighing these options, get in touch with our team and we’ll help you map the right architecture to your data, latency, and compliance needs.

Frequently Asked Questions

What is Jev?

Jev is TypeSafe AI’s first System One model — a non-generative AI model that reads a piece of state plus typed questions and returns typed, calibrated decisions (a choice, a score, or a probability) instead of written text.

What is a System 1 AI model?

A System 1 model is an AI model built to make fast, structured decisions rather than generate text. It answers typed questions against a shared state in a single non-autoregressive pass, returning a value from a predefined schema plus a calibrated confidence score.

Is Jev an LLM?

No. Jev doesn’t generate free-form text and has no chat interface. It shares transformer-based language understanding with LLMs, but its output is restricted to typed values defined by the questions you send it — closer to a very fast, very calibrated classifier than a chatbot.

What is Laya, and how is it different from Jev?

Laya is an open-weights, Apache 2.0-licensed decision model from Convai Innovations, built on a 421-million-parameter ModernBERT encoder. It answers the same three question types as Jev (choice, score, and a yes/no probability called “noul”) but can be self-hosted for free, while Jev is a closed, hosted API priced at $0.042 per million input tokens.

Can Jev or Laya hallucinate?

Neither can return malformed output or a value outside the schema you provide, since there’s no free-text generation step. Both can still confidently select the wrong valid option, especially before fine-tuning and confidence calibration on your own data — so “0% hallucination” refers to output validity, not guaranteed correctness.

What does RLCD stand for?

Reinforcement Learning for Calibrated Decisions, TypeSafe’s term for the training approach behind Jev. It optimizes for confidence scores that are actually trustworthy against strictly proper scoring rules, rather than simply rewarding the model for picking the top answer.

Do I need to fine-tune Laya before using it in production?

For most real workloads, yes. Laya’s zero-shot accuracy on the typed-decisions benchmark (around 0.362) is close to a random baseline; its fine-tuned checkpoint scores considerably higher (0.766) on the same benchmark. Treat the base checkpoints as a fast foundation to specialize on your own labeled data, not a drop-in zero-shot decision engine.

Conclusion

Jev and Laya are early, imperfect entries in a model category that’s likely to keep growing: cheap, fast, typed decisions sitting alongside — not replacing, the generative models most teams already run. The pitch is straightforward even where the benchmark numbers are still contested: most production AI systems make far more small, structured decisions than they generate paragraphs, and paying frontier-LLM prices and frontier-LLM latency for a yes/no answer was never a great trade. Whether that decision layer ends up hosted, self-hosted, or built in-house on a fine-tuned open encoder is exactly the kind of architecture call worth making deliberately — ideally with a partner that has already built the classification, NLP, and agent infrastructure it needs to sit beside.

You might also like: