Blog Network

Firelex · 2026-09-28 · major

Jeff — Jev-compatible 0.8B decision models trained on one home GPU

Jeff is a set of three small open decision models fine-tuned from Qwen3.5 and Gemma 4. They take Jev's request format and return a probability per option in about 22 ms. The 2B scores 83.1% across five benchmarks against Jev's published 83.0%.

GitHub card for the firelex/jeff repository

Three tiny open models that pick between options you describe in plain words, in about 22 milliseconds.

Key specs

Latency (rtx pro 6000)22 ms
Hn points205

Quick facts

MakerFirelex (independent, not affiliated with TypeSafe)
ModelsJeff-Qwen3.5-0.8B, Jeff-Qwen3.5-2B, Jeff-Gemma4-E2B
LicenseCode MIT, weights Apache 2.0
Question typeschoice (up to 255 options), noul, score
TrainingOne RTX PRO 6000; 0.8B in ~2 h, 2B in ~3.5 h
Runs onNVIDIA GPU, CPU, Apple silicon (MLX, Qwen only)

Benchmarks

Overall accuracy, 5 public benchmarks
Jeff-Qwen3.5-2B83.1%
Jeff-Gemma4-E2B81.6%
Jeff-Qwen3.5-0.8B79.1%
Jev (published)83%
AutoJev-27B (published)84.9%
source ↗

What is it?

Jeff is a family of three zero-shot decision models — a 0.8B and a 2B fine-tune of Qwen3.5 and a fine-tune of Gemma 4 E2B — that accept the same request format as TypeSafe's Jev. You describe a situation and list the options in words, and Jeff returns a calibrated probability for each one. The project is independent and says it is not affiliated with or endorsed by TypeSafe.

How does it work?

No text is generated: one forward pass reads a trained answer over the option letters, and a single fitted temperature calibrates the probabilities. The training code starts from Denis Yarats's open AutoJev recipe and adds small student models, a local synthetic-data pipeline with a leak filter, and MLX serving on Apple silicon. All synthetic data was written by the open Qwen3.8-Flash-Next on two DGX Sparks, and training ran on a single RTX PRO 6000 workstation GPU.

Why does it matter?

Routing, moderation and intent calls can run locally in 22–60 ms instead of costing an API call each. The README is clear about the trade-off: Jeff wins on classification and grounding tests such as Financial PhraseBank (96.4%) but stays well below Jev on reasoning-heavy BBH and JudgeBench. A short fine-tune closes much of the gap for a narrow task — a voice-navigation run moved held-out accuracy from 31.7% to 95.8% in under half an hour.

Who is it for?

backend and agent engineers who need fast local classification

Frequently asked questions

Is Jeff made by TypeSafe, the company behind Jev?
No. Jeff uses the same request format as Jev, so code written for Jev can talk to a local Jeff server, but the README states the project is not affiliated with or endorsed by TypeSafe. Jeff's training code started as a fork of Denis Yarats's open-source AutoJev recipe, which is MIT-licensed.
Which Jeff model should I pick, 0.8B or 2B?
For fast option picking the Jeff README recommends the 0.8B. The 2B scores higher on benchmarks (83.1% overall against 79.1%) but the authors found it more cautious and a weaker game player, collecting 41.2 Pac-Man pellets against the 0.8B's 57.0. The 0.8B also runs at 28 ms on an M4 Max against 60 ms for the 2B.
Can Jeff replace a large model for reasoning tasks?
No. The Jeff authors say small models don't reason: expect fast, calibrated choices between options you describe, not multi-step reasoning. On BBH Jeff-Qwen3.5-2B scores 68.0% against Jev's published 94.3%, and on JudgeBench 64.6% against 78.6%. Jeff also handles English text only, and it cannot forecast — describe what each option leads to instead.
Can I fine-tune Jeff on my own data?
Yes. Jeff's weights are Apache 2.0 and the repository includes the AutoJev-based training scripts. The authors fine-tuned it on about 11,000 voice-navigation examples in roughly half an hour on one GPU, moving held-out accuracy from 31.7% to 95.8% at about 40 ms per decision on an M4 Max.

Try it

uv run hf download mstrasser/Jeff-Qwen3.5-0.8B --local-dir checkpoints/jeff-0.8b

Sources · 3 outlets

Tags

  • jeff
  • firelex
  • jev
  • autojev
  • open-weights
  • qwen3-5
  • gemma-4
  • classification
  • zero-shot
  • decision-models
  • calibration
  • mlx
  • small-models

← All releases