Firelex · 2026-09-28 · major
Jeff — Jev-compatible 0.8B decision models trained on one home GPU
Jeff is a set of three small open decision models fine-tuned from Qwen3.5 and Gemma 4. They take Jev's request format and return a probability per option in about 22 ms. The 2B scores 83.1% across five benchmarks against Jev's published 83.0%.
Three tiny open models that pick between options you describe in plain words, in about 22 milliseconds.
Key specs
| Latency (rtx pro 6000) | 22 ms |
|---|---|
| Hn points | 205 |
Quick facts
| Maker | Firelex (independent, not affiliated with TypeSafe) |
|---|---|
| Models | Jeff-Qwen3.5-0.8B, Jeff-Qwen3.5-2B, Jeff-Gemma4-E2B |
| License | Code MIT, weights Apache 2.0 |
| Question types | choice (up to 255 options), noul, score |
| Training | One RTX PRO 6000; 0.8B in ~2 h, 2B in ~3.5 h |
| Runs on | NVIDIA GPU, CPU, Apple silicon (MLX, Qwen only) |
Benchmarks
| Jeff-Qwen3.5-2B | 83.1% | |
|---|---|---|
| Jeff-Gemma4-E2B | 81.6% | |
| Jeff-Qwen3.5-0.8B | 79.1% | |
| Jev (published) | 83% | |
| AutoJev-27B (published) | 84.9% |
What is it?
Jeff is a family of three zero-shot decision models — a 0.8B and a 2B fine-tune of Qwen3.5 and a fine-tune of Gemma 4 E2B — that accept the same request format as TypeSafe's Jev. You describe a situation and list the options in words, and Jeff returns a calibrated probability for each one. The project is independent and says it is not affiliated with or endorsed by TypeSafe.
How does it work?
No text is generated: one forward pass reads a trained answer over the option letters, and a single fitted temperature calibrates the probabilities. The training code starts from Denis Yarats's open AutoJev recipe and adds small student models, a local synthetic-data pipeline with a leak filter, and MLX serving on Apple silicon. All synthetic data was written by the open Qwen3.8-Flash-Next on two DGX Sparks, and training ran on a single RTX PRO 6000 workstation GPU.
Why does it matter?
Routing, moderation and intent calls can run locally in 22–60 ms instead of costing an API call each. The README is clear about the trade-off: Jeff wins on classification and grounding tests such as Financial PhraseBank (96.4%) but stays well below Jev on reasoning-heavy BBH and JudgeBench. A short fine-tune closes much of the gap for a narrow task — a voice-navigation run moved held-out accuracy from 31.7% to 95.8% in under half an hour.
Who is it for?
backend and agent engineers who need fast local classification
Frequently asked questions
- Is Jeff made by TypeSafe, the company behind Jev?
- No. Jeff uses the same request format as Jev, so code written for Jev can talk to a local Jeff server, but the README states the project is not affiliated with or endorsed by TypeSafe. Jeff's training code started as a fork of Denis Yarats's open-source AutoJev recipe, which is MIT-licensed.
- Which Jeff model should I pick, 0.8B or 2B?
- For fast option picking the Jeff README recommends the 0.8B. The 2B scores higher on benchmarks (83.1% overall against 79.1%) but the authors found it more cautious and a weaker game player, collecting 41.2 Pac-Man pellets against the 0.8B's 57.0. The 0.8B also runs at 28 ms on an M4 Max against 60 ms for the 2B.
- Can Jeff replace a large model for reasoning tasks?
- No. The Jeff authors say small models don't reason: expect fast, calibrated choices between options you describe, not multi-step reasoning. On BBH Jeff-Qwen3.5-2B scores 68.0% against Jev's published 94.3%, and on JudgeBench 64.6% against 78.6%. Jeff also handles English text only, and it cannot forecast — describe what each option leads to instead.
- Can I fine-tune Jeff on my own data?
- Yes. Jeff's weights are Apache 2.0 and the repository includes the AutoJev-based training scripts. The authors fine-tuned it on about 11,000 voice-navigation examples in roughly half an hour on one GPU, moving held-out accuracy from 31.7% to 95.8% at about 40 ms per decision on an M4 Max.
Try it
uv run hf download mstrasser/Jeff-Qwen3.5-0.8B --local-dir checkpoints/jeff-0.8b