Newsletter image

Subscribe to the Newsletter

Join 10k+ people to get notified about new posts, news and tips.

Do not worry we don't spam!

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use

Search

GDPR Compliance

We use cookies to ensure you get the best experience on our website. By continuing to use our site, you accept our use of cookies, Privacy Policy, and Terms of Service.

TypeSafe - Inference, Ecosystem

Jev and the open decision models

TypeSafe Jev returns typed decisions instead of text. Eleven open projects now chase it, 7.8x faster on a 2018 GPU. But laya is near chance zero-shot, and that changes the comparison.

TL;DR
  • Typed decisions in one forward pass: choice, score and calibrated probability
  • laya is 7.8x faster than Jev p50 on a 2018 T4, and Apache 2.0
  • But near chance zero-shot; it is a base to fine-tune, not a drop-in swap
System Requirements
RAMRuns on CPU, but the speed case needs a GPU or Apple Silicon
GPUTesla T4 class is enough; 32.8 ms per decision measured there
VRAMSmall: 421M params English, 322M multilingual
✓ Apple Silicon

TypeSafe shipped Jev with a claim that sounds like marketing and mostly is not: a model that never produces a structured output error, because it never produces text. You hand it program state and a set of questions, and it returns typed decisions in a single parallel pass. Two weeks later there are eleven open-source projects chasing it, together holding nearly 89,000 GitHub stars, one of them from Nokia's research arm. The open side is roughly seven times faster on a six-year-old GPU. It also cannot do the one thing Jev is actually sold for, and that is the part worth your attention.

What a typed decision is

Everything else follows from this, so it is worth being precise. An ordinary LLM answers a routing question by generating the word "billing", which you then parse, and which it can get wrong in a hundred formatting ways before it gets the answer wrong. A decision model does not generate. It runs one forward pass and returns a value of a declared type.

Three primitives cover most of it. A choice returns one label from a set you supply, with a probability for every option. A score returns an expected value on an ordinal scale, with the distribution behind it. And a boolean returns a calibrated probability that something is true, between zero and one.

That last word is load-bearing. Laya, the main open implementation, trains these with reinforcement learning against strictly proper scoring rules, which is the technique that makes a reported 0.8 mean something close to eight-in-ten rather than a number the model felt good about. A classifier gives you an argmax. This gives you a distribution you can threshold on.

The practical framing that has stuck in the community is the "semantic if": a branch in your program whose condition is a question about unstructured text, evaluated in milliseconds, returning a typed answer instead of a paragraph you have to interpret.

Jev, the thing being answered

TypeSafe calls Jev the first of its System One models, borrowing Kahneman's fast-intuitive-thinking label. It takes unstructured program state, answers multiple questions against one shared input in a single evaluation pass, and reports a 0% structured output error rate.

PropertyJev
Latency70 to 500 ms end to end, about 100 ms typical
Price$0.042 per million input tokens, output free
Vendor comparison193.6x faster and 444.6x cheaper than GPT-6 Astra and Fable 5.1, on TypeSafe's own workflow evaluations
AvailabilityHosted API, also on OpenRouter
WeightsClosed

Treat the 193.6x as what it is: a vendor number on a vendor benchmark, in the same family as every other claim we have audited this month. The architectural claim underneath it is the durable part. Adding more questions to one call adds token overhead but does not multiply the base latency, because the whole batch is evaluated together. That is a real structural advantage over asking an autoregressive model five questions in sequence.

The open cluster

Nine days produced this. Every repository below exists with working code, and the totals are from the GitHub API on the date at the foot of this article.

ProjectStarsLicenseWhat it does
laya28,881Apache 2.0The open engine: frozen encoder plus a decision head
jev-ultrafast21,477MITbrowser-use's web agent built on the pattern
kev7,981Apache 2.0Trainable decision models on Qwen3.5 and 3.8
fast-jev-compaction7,222MITScores every tool call instead of summarising context
jev-chat-jarvis7,158MITMobile conversation copilot
laya-mlx6,643Apache 2.0Apple Silicon runtime
SemIf-OpenJev4,596MITSemantic ifs on a 3090 at home
ollaya1,002Apache 2.0Local serving behind a TypeSafe-compatible API
AnyJev978Apache 2.0Nokia: turn any LLM into a decision model, no training

Plus two directories, taking the tracked total to 88,826 stars across eleven repositories. There are now 2,137 repositories tagged jev on GitHub. Nokia Applied Research publishing into a nine-day-old ecosystem is the detail that separates this from a weekend hype cycle.

Speed, where open wins outright

Laya's published measurements are on a Tesla T4, a 2018 inference card you can rent for pennies or find in any older server.

SetupLatencyHardware
Jev236 to 276 ms p50Hosted
laya, English39.5 msTesla T4
laya, multilingual32.8 msTesla T4
laya, 50 questions batched6.8 ms per questionTesla T4
laya-mlx7 to 14 msM3 Max
laya, CPU or MPS193 to 464 msCPU

Laya claims 7.8x faster than Jev's published p50, and the arithmetic holds. The checkpoints are small enough that this is unsurprising rather than miraculous: 421M parameters on ModernBERT-large for English, 322M on mmBERT-base for the multilingual build covering 100-plus languages. A router picks between them by detecting script and language in under a millisecond.

Note the last row. On CPU you are back to Jev's latency band. The speed story needs a GPU or Apple Silicon, and on a plain server CPU the hosted API is competitive.

The catch that reframes everything

Laya's own README states it without being asked: "Laya is a fast base to specialise, not a zero-shot decision engine."

The numbers behind that sentence are stark. Zero-shot, the base checkpoints score 0.362 and 0.352 on typed decisions against a 0.318 random baseline. That is near chance. After fine-tuning on the task, 0.766.

Jev's entire proposition is zero-shot. You write questions, you get answers, you train nothing. So laya is not a drop-in replacement, it is a fast foundation that requires you to bring labelled data and a training run. For a team that already has that data, it is an excellent trade: you get 7x the speed, Apache 2.0 weights and no per-token bill. For a team that wanted to write four questions on a Tuesday afternoon, it is a different project entirely.

Credit where it is due. Publishing a near-chance zero-shot number in your own README, next to the benchmark you do win, is the behaviour we keep asking vendors for. Laya also documents that both checkpoints ship over-confident and need domain temperature fitting before you threshold on the probabilities, and that the multilingual build has a position bias on score questions.

The two projects that actually close the gap

Which is why the small repositories at the bottom of the table are the interesting ones.

AnyJev, from Nokia's applied research group, turns any LLM into a Jev-style decision model with typed decisions and real probabilities and no training. That is the zero-shot property laya lacks, reached from the opposite direction: instead of a small specialised encoder, constrain a general model you already run.

Ollaya describes itself as Ollama for decision models. It pulls and serves laya, NLI models and GLiClass locally behind a TypeSafe-compatible API, which makes it the literal drop-in swap: point your existing Jev client at localhost and change nothing else.

Between them they cover the two things that stop a team adopting the open side, which is no training data and no migration path. Neither has a thousand stars yet. Both are more strategically important than the projects above them.

Trying it

Laya installs from PyPI and needs Python 3.10 or newer. Extras exist for HTTP serving, an MCP server and LangChain.

python -m pip install laya

The API is a router plus a dictionary of questions, each with a declared type:

from laya import Router

router = Router()
result = router.predict(
    {"body": "I was charged twice. Refund or we cancel."},
    {
        "department": {
            "type": "choice",
            "instructions": "Which department handles this?",
            "criteria": {"billing": "refunds", "technical": "bugs"},
        }
    },
)
print(result["answers"]["department"]["choice"])

There is a CLI too, which is the fastest way to get a feel for whether the shape fits your problem before you commit to a fine-tune.

laya "I was charged twice" --predict
laya --batch tickets.txt --predict --json

Where this fits

Use a decision model where you are currently asking an LLM a question whose answer is a label, a number or a boolean, and then parsing the reply. Ticket routing, content moderation, tool selection in an agent loop, retrieval reranking, guardrails. Anywhere you wrote a regex against model output and felt bad about it.

Do not use one where you need the text. This is not a small language model and it will not write anything. The interesting agent application is the one fast-jev-compaction demonstrates: rather than asking a big model to summarise a context window, score every tool call and result with a small fast model and keep what scores high. That is a genuinely different use of the primitive.

If you are evaluating today: try ollaya first if you already pay for Jev and want to know whether local serving is viable, try AnyJev if you want typed decisions out of a model you already run, and go to laya when you have labelled data and latency is the constraint that matters.

Sources and further reading

Ten minutes: pip install laya, run the CLI against twenty lines of your own support tickets or logs, and look at the probabilities rather than the labels. If the confident answers are right and the uncertain ones are the genuinely ambiguous ones, the primitive fits your problem and the only remaining question is whether you fine-tune or reach for AnyJev. If the probabilities look like noise, you have learned that in ten minutes for nothing.

Tested on: not independently tested. Jev's latency, pricing and comparison figures are TypeSafe's own published numbers. Laya's latency, accuracy and calibration figures, including the near-chance zero-shot baseline, are from its own README and measured on the author's Tesla T4. No independent replication of either side exists at the time of writing. Star counts, fork counts and topic totals were read from the GitHub API on the date below and are moving fast: this cluster is nine days old.
Date checked: 2026-09-30

Prev Article
ZCode and the harness wars
Next Article
Herdr

Related to this topic: