Skip to main content
← /writing
  • #ai-economics
  • #genai
  • #ai-strategy

Jev Is Not a Faster LLM

TypeSafe's Jev gives up prose for typed, probabilistic decisions. The engineering work is defining those decisions, measuring their errors, and deciding when software should ask for help.

Vinny Carpenter13 min read2.6k words

never the stack · audio edition

Jev Is Not a Faster LLM

17:17

In 1865, the English economist William Stanley Jevons published The Coal Question. James Watt's steam engine burned far less coal per unit of work than the Newcomen engine before it, and England's coal consumption rose anyway. Efficiency made steam power cheaper, cheaper steam power was worth using in more places, and total demand outran the savings. "It is wholly a confusion of ideas to suppose that the economical use of fuel is equivalent to a diminished consumption," Jevons wrote. "The very contrary is the truth."

Economists call this the Jevons paradox, and I used it in August to explain where the saved hour goes when AI makes work faster. But Jevons made a smaller point I find more useful. Thomas Savery's earlier engine was so inefficient that the cost of running it kept it out of use. It consumed no coal, he noted, because it consumed too much.

On Sept. 15, TypeSafe AI released Jev, a model named after Jevons. The company expects cheaper machine intelligence to open many more uses. Jev gives up something most new models compete to improve. It doesn't write. You hand it a situation and typed questions, and it returns structured decisions and probabilities.

I'm building my first workflow on it, a triage pipeline for the feedback my apps collect. I want a development queue that surfaces actionable bugs and requests without making me read every item closely. Jev is in early access, the performance numbers below belong to TypeSafe, and I don't have results of my own yet.

The interesting part is the contract: define the judgment, measure its errors, and write down when the system should ask for help.

Cheaper decisions mean more decisions

In its launch post, TypeSafe lists Jev at $0.042 per million input tokens with no output charge. It reports response times of 70 to 500 milliseconds. At that input price, a request with 1,000 tokens costs $0.000042. That's cheap enough to consider decisions you might otherwise skip. Whether it returns fast enough for a particular application is a separate question, and one I'd measure from where that application runs.

Some judgment calls go unmade for Savery's reason. At enough volume, asking a model whether a log line matters can cost too much in money, latency, or operational effort. Teams sample, batch, or skip. Those decisions consume no tokens because each one would consume too many.

Jevons offers an analogy, and software demand doesn't have to follow it. Lower costs could mean we check more logs and comments, then find judgments we never considered automating. Whether that happens depends on whether the decisions are useful and whether the rest of the workflow can absorb them.

For a leader, that second-order effect matters more than the API price. Each automated judgment can influence what gets attention and what gets ignored. As those judgments multiply, so does the work of deciding which errors are acceptable.

Cheap decisions can also create expensive attention demands. A system that flags everything has saved me little. I want fewer, better priorities. As implementation gets cheaper, the constraint moves toward evaluation and judgment. Cheap decisions could move it there faster.

System 1 answers fast and calls for help

TypeSafe calls Jev a System One Model, borrowing from Daniel Kahneman's Thinking, Fast and Slow. System 1 is fast and automatic. System 2 is slow and deliberate, the mode you engage to multiply 17 by 24. System 1 runs all the time and handles nearly everything, and it calls on System 2 when it meets something it can't settle.

It's a useful architectural analogy. Some judgments can be made quickly with the right context; others need sustained reasoning. It isn't an account of how these models think internally. TypeSafe's documentation asks you to keep each question small enough that a knowledgeable person could answer it in a few seconds.

The warning matters as much as the speed. A bat and a ball cost $1.10 together, and the bat costs $1 more than the ball. The tempting answer is that the ball costs 10 cents. The right answer is 5 cents. A fast answer can feel certain and still be wrong.

That's the property I'd test before trusting a decision model. When it assigns an outcome a probability of 95%, does that outcome occur about 95% of the time across comparable cases? And does its uncertainty signal help identify the cases where it needs help?

A model built around decisions

It's tempting to file Jev under small, fast LLMs. But TypeSafe is proposing a specialization. It gives up general text generation and builds the model around typed, probabilistic decisions.

That distinction needs a qualification. LLMs already produce structured outputs. APIs can enforce supported schemas and return parsed values that application code can use directly. OpenAI's Structured Outputs does this, for example. A program branching on a model's answer isn't new.

The question is whether a model designed for that job beats a small LLM answering the same questions under the same schema. I'd compare them on probabilities, latency, and cost. That's the comparison I want to make on my own workload.

Jev exposes three primitives. Choice selects an option from a list. Score rates the input against a rubric. Noul evaluates a true-or-false statement and returns a value between 0 and 1. Choice and Score return probability distributions and a separate confidence field. Questions in a request are evaluated in parallel and independently against the same state.

A general text-generating LLMJev
OutputProse, code, or schema-constrained valuesChoice, Score, and Noul results
GeneratesOne token at a timeAll answers in parallel
Trained towardResponses people prefer, or outputs that pass a checkCalibrated probabilities
Role in this workflowA baseline classifier; later, diagnosis and draftingClassification, scoring, and probabilistic checks

The generation and training rows are TypeSafe's description, and I haven't verified either. I need to test both models on accuracy, uncertainty, latency, and total cost, using the same feedback. Both can fail the same way, with a schema-valid answer that's wrong.

Jev can't write a reply to a reviewer, draft a fix, or explain its answer in a paragraph. TypeSafe advises decomposing complex judgments into small questions and combining the answers in code.

That narrower output contract also explains why I'd read the company's claim that Jev "can't hallucinate" carefully. TypeSafe guarantees schema matching; its own notes say the zero in its chart comes from that guarantee rather than an empirical measurement. A well-typed answer can still be wrong. A crash report can be labeled praise, and the type system won't object.

TypeSafe does publish useful caveats alongside its claims. Its workflow evaluations use other models' predictions as references rather than human ground truth, and its own team constructed the workflows. I'd like every vendor to write launch posts this way. The results give me a reason to investigate, and I still have to test against the decisions I need to make.

I'm starting with my app feedback

I actively track feedback on 10 apps across Apple's App Store, Google Play, and the web. Feature requests alone usually run a couple of dozen a month. They arrive mixed with bug reports, usability complaints, pricing comments, and praise. Some of these apps have been in the App Store for years, so I also have a history to work through. It includes bug reports, feature requests, and crash data.

A tiny difference in token price won't make or break this workload. It's a bounded place to test the contract before testing the scale. I know the apps, I can label the feedback, and I can inspect the mistakes. A small LLM might already do the job well enough. Jev has to earn its place.

I'm starting with three questions and an escalation rule.

The questionTypeWhat the code does with the answer
What is the primary issue: a bug, a request, a usability problem, a pricing complaint, or praise?ChoiceSuggests the initial queue
How severe is the reported problem under my rubric?ScoreHelps prioritize problem reports
Does the feedback describe something specific I can act on?NoulHelps distinguish concrete reports from vague reactions
Should I review this before trusting the routing?Policy using uncertainty and consequencesEscalates ambiguous or consequential cases

The clean examples are easy. A specific crash report belongs in the bug queue. A request for an export option belongs in the feature queue. When diagnosis or a reply is needed, that work goes to an LLM.

The harder cases are mixed. Consider this review, which I made up to illustrate the problem:

I love this app and use it every day, but yesterday it lost my notes after syncing. Also, could you add PDF export?

One Choice can select a primary queue, but it can't preserve every issue in that review. If I route it as praise, I've buried possible data loss. If I route it as a bug, the feature request can disappear. A confident answer won't fix an inadequate taxonomy.

That may mean asking separate questions about bugs and requests, allowing more than one queue, or sending mixed reports to me. The model evaluates each question independently; the code has to reconcile the answers. The questions and categories can be wrong even when the model answers them correctly.

What the probabilities have to earn

I need to distinguish two things the API returns. The probabilities describe the possible outcomes. TypeSafe's confidence field summarizes the shape of that distribution. The documentation doesn't establish that a confidence value of 0.95 means a 95% chance of correctness. Noul has no separate confidence field.

I'll label a sample of historical feedback by hand and compare the outcome probabilities with observed results. For classifications assigned roughly 90% probability, do they turn out to be correct roughly 90% of the time? I'll also test whether the confidence field is useful for choosing escalation thresholds. Those are related tests, but they aren't interchangeable.

Even good calibration won't settle whether the queue is useful. A model can perform well overall and still miss the reports I care about most. Missing a data-loss bug costs more than filing praise in the wrong place. That means thresholds should depend on the consequences, and suspected data loss should get my attention even when the model seems certain.

I'll compare Jev with a small LLM producing the same schema and report three practical outcomes. The first is how many actionable bugs each one misses. The second is how much feedback still needs my review, and the third is how much time I save after correcting mistakes. I'll inspect a sample of the automatically routed items too. Reviewing only uncertain answers would hide the confident errors.

The payoff I'm looking for is less reading without losing important reports. A cheaper call that produces a worse queue hasn't improved the workflow.

Route by the kind of answer you need

The review pipeline suggests a larger design, with several lanes for different kinds of work.

  • Defined calculations, validation, and explicit rules go to ordinary code.
  • Straightforward summaries or drafts can go to a small model that meets the quality requirement.
  • Diagnosis and difficult reasoning go to a more capable model, with extended reasoning when it helps.
  • Typed judgments can go to a decision model, if it performs well on the workload.

Often the caller already knows the lane. A summary request doesn't need another model to discover that it needs writing. A defined calculation doesn't need an AI router. Route those requests directly in code.

A learned router becomes useful when the required work is ambiguous. I've been calling that piece a decision engine, for lack of a better name. It might decide whether a report needs classification, diagnosis, or clarification. Its routing judgment is one more decision to evaluate, and it doesn't justify putting a model in front of every call.

A routing flow. When the required work is already known, explicit routing rules send the request to its lane. When it isn't, a decision engine evaluates the routing judgment and a routing policy weighs uncertainty and consequences. An accepted route goes to the routing rules. A request that needs clarification goes to a capable model or a person first. The four lanes are ordinary code, a small model, a capable model, and a decision model.

Escalating to resolve the route is different from choosing the reasoning lane. The first asks what work is needed; the second does that work. A routing policy should combine measured uncertainty, the cost of mistakes, and explicit constraints on what can happen automatically.

Model routers already exist, and I made a routing argument for local models in August. Jev adds a candidate specialized for decisions. Whether it should also make routing decisions is a hypothesis to test.

This borrows Kahneman's division of labor, and inherits his warning. The bat-and-ball failure happens when the fast answer is wrong and never asks for help. A confident router can make that mistake too. Its predictions should inform the policy; they shouldn't be the whole policy.

What would change my mind

Three findings would weaken the case for Jev in my workflow. First, outcome probabilities that don't hold up on my feedback. Second, decisions that change materially when I reword the same facts. Third, no useful improvement over a small LLM once I include missed bugs, review time, latency, and cost.

Any of those could make me keep the workflow and change the model. If automatic triage costs more attention than it saves, I'd reconsider the workflow itself.

Competition is a different question. TypeSafe is a young company with an early-access model and a proprietary API. If another vendor offers a better decision endpoint, that changes my supplier rather than disproving the pattern. I'll keep the questions, rubrics, and thresholds in my code, with the vendor behind an interface I control. It's the same argument I made in The Agent Is Not the Platform.

Choose which decisions should be cheap

If you lead an engineering team, the first step doesn't require Jev or a waitlist. Pick one workflow where a model makes a judgment that influences what happens next. Write down the question, the allowed answers, and the cost of getting each one wrong. Keep a record of what the model predicted and what a person would have decided.

That record lets you compare the model you already run with whatever comes next. It also forces a harder conversation about which decisions should happen automatically and which deserve someone's attention regardless of price.

Jevons gives me a reason to expect cheaper judgment to find more uses. My app-feedback pipeline gives me a way to test whether one of those uses is worth having. I'll publish the missed bugs, the review burden, and the time saved. That's what the probabilities have to earn.

// found this useful? share it

Post on X Share to LinkedIn
Vinny Carpenter

Written by Vinny Carpenter

VP Engineering · 30+ years building software

I lead engineering teams building cloud-native platforms at a Fortune 100 company. I write about engineering leadership, AI-assisted development, platform strategy, and the hard lessons that come from shipping at scale.

keep reading