Newsletter image

Subscribe to the Newsletter

Join 10k+ people to get notified about new posts, news and tips.

Do not worry we don't spam!

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use

Search

GDPR Compliance

We use cookies to ensure you get the best experience on our website. By continuing to use our site, you accept our use of cookies, Privacy Policy, and Terms of Service.

Jeff - Inference, Language Model

Jeff

Jeff fine-tunes Qwen3.5 and Gemma 4 into zero-shot decision models scoring 82.0% against Jev's 83.0%, at about 22 ms per decision on local hardware.

TL;DR
  • Three checkpoints from 1.7 GB to 9.3 GB, fine-tuned from Qwen3.5 and Gemma 4
  • 82.0% zero-shot across 4,599 questions against Jev's 83.0%
  • About 22 ms per decision on an RTX PRO 6000, 28 ms on an M4 Max

Two days ago we wrote that the open answer to TypeSafe's Jev had one hole in it. Laya is fast and Apache 2.0 and near chance zero-shot, and Jev's whole proposition is zero-shot, so the open side was a fast base you had to train rather than a drop-in replacement. Jeff was published the same week and closes most of that gap: 82.0% against Jev's 83.0% across 4,599 questions, no training required, running in about 22 ms on a local GPU. This is part two of our decision models series, and it arrived faster than the argument did.

What Jeff is

Jeff is three fine-tuned checkpoints that take the same request shape Jev does and return calibrated probabilities from a single forward pass. You describe a situation, list the options in plain language, and get a distribution back. No training on your categories, which is the part that matters.

CheckpointBaseSize on disk
Jeff-Qwen3.5-0.8BQwen3.51.7 GB
Jeff-Qwen3.5-2BQwen3.54.2 GB
Jeff-Gemma4-E2BGemma 49.3 GB

The training recipe is refreshingly plain: full-weight fine-tuning, one epoch, batches of 256, cross-entropy over the option letters, then a single fitted temperature for calibration. Synthetic training data was generated with Qwen3.8-Flash-Next on two DGX Sparks, and the project states no proprietary model output was included, which matters if you intend to ship the weights.

Code is MIT, weights are Apache 2.0, and training data sources keep their original licenses. The project is explicit that it is "not affiliated with or endorsed by TypeSafe, the makers of Jev", and that it starts from the open AutoJev recipe while using the same request format.

The number that closes the gap

Across 4,599 questions, these are the project's own reported results.

ModelOverallGap to Jev
Jev (TypeSafe)83.0%baseline
Jeff-Qwen3.5-2B82.0%1.0 point
Jeff-Gemma4-E2B81.6%1.4 points
Jeff-Qwen3.5-0.8B79.1%3.9 points

One point behind a hosted commercial model, from a 4.2 GB open checkpoint you run yourself. Compare that to where laya sits zero-shot, at 0.362 against a 0.318 random baseline, and the shape of the open side changes completely in the space of a week.

The per-task detail is more useful than the average. Financial PhraseBank lands at 94.7 to 96.1% and RAGTruth at 85.6 to 87.7%, both strong. Reasoning-heavy suites like BBH and JudgeBench are where it falls away, which is exactly what you would predict from a 0.8B to 2B model with no chain of thought. The project says so directly: "small models don't reason."

Speed, and where it runs

About 22 ms per decision on an RTX PRO 6000, and 28 ms on an Apple M4 Max through MLX. Jeff cites Jev's published figure as 114 to 212 ms per call.

Worth flagging: we have now seen three different latency figures for Jev from three sources, the vendor's own 70 to 500 ms band, laya's cited 236 to 276 ms p50, and Jeff's 114 to 212 ms. They are probably measuring different things, end to end versus model time, and none of them is independently verified. Use the ratio rather than the absolute, and the ratio is consistently in the open side's favour once you own the hardware.

The Apple Silicon number is the one that should interest most readers. 28 ms on an M4 Max means the whole category runs on a laptop, which is not true of anything else we have covered this month.

Running it

uv sync --no-default-groups --extra cuda
uv run --no-default-groups hf download mstrasser/Jeff-Qwen3.5-0.8B --local-dir checkpoints/jeff-0.8b
JEFF_CHECKPOINT=checkpoints/jeff-0.8b PORT=8765 uv run --no-default-groups jeff-serve

The request format is the same shape the rest of this category uses, which is the quiet advantage of everyone copying Jev's interface: swapping implementations is a base URL change.

{
  "model": "jeff-latest",
  "state": "situation description",
  "questions": {
    "route": {
      "type": "choice",
      "instructions": "Which team handles this?",
      "criteria": {"1": "billing", "2": "technical"}
    }
  }
}

The limitations, which are unusually well documented

Option counts are capped: the Qwen checkpoints accept up to 254 options, Gemma4-E2B stops at 26. It is English and text only. Prompt wording moves the results noticeably, which is true of the whole category and rarely admitted.

Then there is the finding the project did not have to publish. On its games evaluation, Pac-Man, Frogger and Doom, the smaller 0.8B model plays better than the 2B despite scoring lower on the benchmark average. A model that wins on aggregate accuracy and loses on an interactive task is a useful warning about what a single headline number is worth, and it is to the authors' credit that it is in the README rather than left for someone else to find.

Which of these to reach for

If you haveUseBecause
Labelled data and latency pressurelayaFastest once specialised, 6.8 ms per question batched
No labelled dataJeff82.0% zero-shot, one point off the commercial product
An existing Jev deploymentollayaTypeSafe-compatible API, swap the base URL
An LLM you already runAnyJevTyped decisions from a general model, no new weights

A week ago this table had a hole in the second row. That is how fast this particular corner is moving, and it is the reason the date at the foot of this article is doing more work than usual.

Sources and further reading

Ten minutes: download the 0.8B checkpoint at 1.7 GB, serve it, and post the same routing question you would send to Jev. If the probabilities look sane on your own data, you have replaced a metered API with a file on your disk, and the only thing you gave up is one point of benchmark average.

Tested on: not independently tested. Accuracy, latency, training recipe and limitation figures are all reported by the Jeff project in its own README, including the Jev comparison baseline. The three conflicting published latency figures for Jev are noted in the text; none is independently verified. Repository statistics were read from the GitHub API on the date below, when Jeff was two days old and stood at 1,128 stars.
Date checked: 2026-09-30

Prev Article
Hermes Agent
Next Article
anydoc

Related to this topic: