Another interesting week in AI...
Gemini 4 Argon is pushing coding and enterprise workflows forward. Meanwhile, Anthropic’s Opus 5.5 announcement reports an early tester’s single-prompt game comparison where graphics and polish stood out. The demos are getting wild.
But a quieter architecture question has my attention.
For the past few years, we have been putting LLMs everywhere.
Generate an email? LLM.
Summarize a contract? LLM.
Write code? LLM.
Need to decide whether a customer should receive a refund? Also LLM.
That last one is where things get interesting.
Generating an answer and making a decision are not necessarily the same problem.
This week in my diary, I have been looking at Jev, TypeSafe’s System One model. I think the bigger story is not just Jev. It is the idea of decision models as a distinct primitive inside software.
Instead of generating open-ended text, Jev answers bounded questions with typed, probabilistic results. There is no paragraph to parse, but that does not mean the answer is correct or the application should automatically act.
The surrounding software still owns the workflow.
If you learn better visually, check out my Visual AI Library on Gumroad. I use these cheat sheets myself before customer meetings to refresh key concepts and explain them more clearly. There is a lot of noise in AI. This library gives you the distilled version, so you can quickly review important concepts before interviews, exams, presentations, or customer conversations.
1. So what exactly is Jev?
TypeSafe introduced Jev in a post dated September 15, 2026. It describes the interface simply: supply state and typed questions, receive structured answers your code can use.
Traditional LLMs generate sequences of tokens. Jev takes a different approach. It gives up open-ended string generation and focuses on decisions among defined possibilities.
TypeSafe calls its training approach Reinforcement Learning for Calibrated Decisions (RLCD). That is a provider-described method, not proof that every domain already has perfectly calibrated predictions.
The simplest way I think about the distinction is:
LLMs generate. Decision models judge. Code decides what happens next.
For example, an LLM might write:
“Based on the customer’s cancellation date and billing history, I recommend a refund.”
A typed decision model approaches the same problem differently. Instead of generating freeform text, it returns a structured result such as a class, score, or probability that the application can use directly.
So Jev could answer a narrower question:
“Does this request fit the supplied refund policy?”
A Noul result might be 0.96, representing the probability of yes.
The application can then combine that probability with its own business rules to decide what happens next.
Jev currently handles text, including text inside JSON, rather than images, audio, or video. It also does not generate a reasoning explanation.
And importantly, modern LLMs can also return structured outputs. So the real question is not whether LLMs can do this.
It is whether a model specialized for narrow decisions can do it faster, cheaper, or more consistently than a general-purpose model.
2. Who actually asks Jev the question?
The customer is not asking Jev, “Should you refund me?”
Your application asks Jev.
Imagine Maya, an engineer building a support assistant. A customer writes: “I cancelled last week, but you charged me again. Please refund this.”
Maya’s application gathers permitted conversation history, account facts, recent charges, refund history and the relevant policy. Dates and charge calculations come from code and authoritative systems, not a model’s guess.
It sends the relevant state alongside predefined questions:
Intent: Refund, cancel, change plan or something else?
Policy fit: Does this request match the supplied refund policy?
Review: Does the message contain a reason to involve a person?
These can be independent questions about the same state. Their answers do not secretly become instructions for one another.
The application interprets the results. A strong policy-fit signal can propose a refund for backend checks; an unclear result can request evidence or review.
But eligibility, permissions, payment limits and required approvals still apply before any payout. A probability threshold cannot bypass them.
Now we have separated two responsibilities: Jev provides semantic judgment. Code controls what happens because of it.
3. Jev has three basic decision primitives
The product names are unfamiliar. The jobs are not: check, choose, score.
Noul returns a number between zero and one: the probability of yes. Your application decides how to interpret it.
Choice returns the selected option and probabilities across the choices. Score returns a numeric result for your ordered rubric, which can be fractional.
Choice and Score also provide confidence. Noul does not provide a separate confidence field.
Production workflows rarely depend on one giant decision. They depend on many small judgments that code combines.
4. Think of it as a smart if-statement
Traditional software might contain this illustrative Python:
if invoice.total_usd > 10_000:
require_approval(invoice)
There is no reason to use AI to compare those numbers.
But consider: “If this invoice description appears inconsistent with the work we bought, route it for review.”
How do you express that as a simple string comparison?
“Cloud infrastructure consulting” and “Azure architecture advisory services” could describe the same work. Different wording does not necessarily mean a mismatch.
This is where a bounded semantic question helps. The model estimates whether the descriptions match; code uses that signal to continue checking or flag the invoice.
The signal is not payment approval.
I find TypeSafe’s “smart if-statements” framing useful: code handles exact rules, decision models handle bounded interpretation, and reasoning models handle open-ended complexity.
5. System One versus System Two AI
The name borrows Daniel Kahneman’s System 1 and System 2 framing in Thinking, Fast and Slow: quick judgments versus deliberate reasoning.
For AI, I treat this as a task analogy, not a claim about biological thinking or a guarantee that every LLM is slow.
Imagine asking your most senior architect hundreds of thousands of times: “Billing or technical support?”
They can answer. That does not mean they are the right resource for every repetition.
The same principle could apply to models. A reasoning model can plan and write, a decision model can handle known-option judgments, tools can perform scoped work, and people can handle consequential uncertainty.
That is a division of responsibilities, not a mandatory chain every request must traverse.
6. Where decision models fit in the enterprise
TypeSafe publishes four enterprise-style workflow evaluations, not production customer case studies. Here are simplified examples of Jev’s role in each
Customer service
A customer writes: “I cancelled last week, but you charged me again. Please refund this.”
The application already has the customer message, account details, cancellation date, billing history, and refund policy. Instead of asking Jev to write a response, it asks a narrow question:
“What does the customer want?”
Possible choices might be refund, cancel, change plan, or other.
Jev returns refund with a high confidence score.
That does not trigger an immediate payment. The application opens the refund workflow, where code checks the charge, eligibility, permissions, and any required approvals.
Jev interprets the intent. Code decides whether the refund is allowed. A language model can still generate the customer-facing reply.
Invoice processing
An invoice says “Azure architecture advisory.” The purchase order says “cloud consulting.”
The wording is different, but the underlying work may be the same.
The application can ask Jev:
“Do these descriptions refer to the same work?”
Jev might return a Noul probability such as 0.93 for yes.
That signal becomes one input into the invoice workflow. The application can combine it with deterministic checks for amounts, dates, vendor identity, duplicate invoices, contract terms, and approval requirements.
Jev handles the semantic comparison. It does not calculate totals or authorize payment.
Security incident response
An alert shows a server login followed by customer-data downloads. The maintenance record only authorizes a software update.
The application asks:
“Does this activity match the approved maintenance work?”
Possible choices might be matches, does not match, or unclear.
If Jev returns does_not_match, the application can raise the severity or route the incident for investigation.
That still does not mean Jev can shut down the server. Containment actions remain governed by predefined security policies, permissions, and approval rules.
Jev helps interpret the activity. Code determines what response is allowed.
Agent trace observability
A customer asks an agent to cancel a subscription.
The agent gives a helpful explanation of how to cancel, but never actually performs the cancellation.
The response sounds good, but the task failed.
The application can give Jev the customer request, conversation, tool calls, tool results, and final response, then ask:
“Was the customer’s task completed?”
Possible choices might be completed, not completed, or unclear.
If Jev returns not_completed, the application can flag the trace for review instead of treating a fluent answer as success.
This is an important distinction for agent evaluation. The question is no longer only whether the response was good. It is whether the system actually completed the task.
Across all four examples, the pattern is the same:
Context → bounded question → structured judgment → business rules → action
Jev handles the semantic judgment. The surrounding application decides what happens next.
AI judges. Code governs. Software acts.
7. Why probabilities matter
Suppose a Noul question about policy fit returns 0.51. Now compare that with 0.995.
If code turns both into “yes,” an important distinction disappears.
An uncertain result might justify more evidence, a reasoning model or human review. A stronger signal might justify less semantic investigation, not fewer mandatory controls.
Do not confuse a probability with confidence. Choice and Score provide a separate confidence statistic derived from their distributions. Noul provides its yes-probability without that extra field.
TypeSafe describes calibrated probabilities, but calibration concerns outcomes across groups of predictions. It does not guarantee one decision is correct.
There is no universal 95% or 98% threshold that makes a payment, security response or medical decision safe. Test thresholds on your data, with the cost of errors and actual review capacity in mind.
Autonomy is an application policy, not simply a property of how intelligent the model is.
8. Does typed output mean Jev cannot be wrong?
No.
Predefined outputs address a structural problem: the model is not supposed to invent a category outside the schema.
They do not solve every judgment problem.
Separate two questions:
Structural reliability: Does the answer fit the expected type?
Decision accuracy: Did it choose the right answer?
A typed wrong answer is still a wrong answer.
Malicious text can also influence Jev’s answers. Keep customer text and retrieved content on the untrusted side of the boundary. A “no attack detected” label does not make them safe instructions.
You still need evaluations, calibration testing, auditability, escalation rules and human review where risk requires it.
9. What about the speed and cost claims?
TypeSafe mentions an end-to-end response time of 70ms-500ms. This can range from 40x-200x faster for the same levels of frontier intelligence for System One shaped queries.
Test your workflow against the incumbent route, simple rules and a small structured-output LLM. Measure accuracy, calibration, end-to-end latency, cost, escalation workload, false positives, false negatives and the business outcome.
Vendor model latency does not include every retrieval step, retry, review or correction. Cheaper inference is not automatically cheaper operations.
The bigger question is not whether your improvement matches a headline multiplier.
Why spend frontier-model intelligence on every tiny decision when a narrower approach might do the job?
10. This is an intelligence-routing problem
I recently wrote about token maxing versus value maxing. The principle applies here too.
The goal is not merely “use fewer tokens.” It is to use the least expensive approach that reliably produces the required outcome.
These are resources to route between, not an escalation ladder where permission checks appear only at the end. Code-owned controls apply on every action path.
The mistake is sending every problem straight to the largest model.
11. Where I would test Jev first
I would start with a reversible, well-labeled task: intent classification, ticket routing, document categorization, retrieval relevance or trace-review prioritization.
Keep contradictory retrieval evidence when it matters; relevance filtering should not erase inconvenient facts.
Be more cautious about moving money, blocking critical systems, making medical decisions or taking legally consequential actions. In those workflows, a model judgment can be one signal, never the entire authorization mechanism.
Learn by doing! Here are some awesome list of demos and repositories on Github.
The bottom line
Pick one repeated judgment with known possible answers, test a decision model against your existing approach, and keep permissions, exact rules and required approvals in your software.
The smartest architecture may not use the smartest model everywhere. It may know exactly where intelligence is actually required.
♻️ If this was useful, share it with someone building with AI.
✉️ Subscribe at newsletter.karuparti.com so you never miss an edition.
Anu Karuparti Creator, Diary of an AI Architect How enterprises actually ship AI to production
Read by 3,000+ FDEs, AI Architects, and Engineering Leaders from Microsoft, Google, IBM, PwC and others.
Connect with me on LinkedIn.
Want to partner? Email me at anurag.karuparti@gmail.com.
References
Google: Gemini 4 Argon. Coding and enterprise-workflow framing.
Anthropic: Claude Opus 5.5. Attributed early tester game comparison, not an independent graphics benchmark.
Anthropic: Structured outputs. A fair generative-model baseline.
My previous edition: 10 ways to stop wasting tokens. Value-maxing framework and the Jev teaser.
TypeSafe: Introducing System One models and Jev. September 15 article dateline, RLCD, workflow benchmark claims, reference-model methodology and stated limitations.
TypeSafe: Introduction and System One. Model scope, typed decisions, independent questions and task analogy.
TypeSafe: Quick start and Score. Primitive response shapes and fractional ordered scores.
TypeSafe: Confidence. Distribution-derived confidence, Noul distinction and workflow-specific thresholds.
TypeSafe: How to build with System One. Code-owned workflows, narrow questions, policy comparison and review signals.
TypeSafe workflow evaluations: customer service, invoice processing, security incidents, and agent trace observability. Internally authored evaluation scenarios, not claimed customer deployments.
TypeSafe: Jev 1.13 jaggedness. Exact arithmetic/date and adversarial-state limitations.
TypeSafe: Classifying RAG passages. Relevance, evidence, contradictions and screening as separate signals.
TypeSafe: Models and Legal. Version pinning, request limits and applicable provider terms.
Disclaimer: The stories and scenarios in this article are hypothetical, inspired by patterns observed across similar real-world experiences. They are used to convey key concepts more effectively and do not represent any specific individual or organization.














"Code decides what happens next" is a good line to hand anyone designing agents. I agree a probability threshold shouldn't override permissions or approval limits. Where I'd push a little: the threshold is a policy decision too, so someone has to own it and revisit it when the model changes. Otherwise a 0.8 cut-off quietly becomes a rule nobody signed off on. I run JevMade, a directory of Jev guides and experiments, and this is the kind of write-up I'd point people to.