Jeff is a set of small, fast decision models fine-tuned from Qwen3.5 and Gemma 4 for zero-shot classification. They slot into your code with the same request format as Jev, returning calibrated probabilities for each option in a single forward pass—about 22 ms per decision on an RTX PRO 6000 and 28 ms on an Apple M4 Max (MLX). No generated text or parsing is needed.
Zero-shot means the options can be anything: support queues, user intents, moderation labels, voice commands, or game moves. Your categories don't need to appear in the training data; you describe them, and Jeff picks. These are very small models that make extremely fast, well-calibrated judgement calls. On benchmarks they approach, and sometimes beat, Jev; but at this size their reasoning won't match Jev's larger model. If zero-shot accuracy isn't enough, a short fine-tune on your own examples can improve accuracy dramatically—a voice-navigation fine-tune moved held-out accuracy from 31.7% to 95.8% in under half an hour on one GPU.
Training is done entirely on local hardware: one RTX PRO 6000 workstation GPU (the 0.8B model trains in about 2 hours, the 2B in about 3.5), with all synthetic training data written by an open model (Qwen3.8-Flash-Next) on two DGX Sparks, and testing on a Mac Book. No cloud GPUs or closed-model output were used in training data; a closed model only spot-checked synthetic data quality. Jeff uses the same request format as Jev but is not affiliated with or endorsed by Type Safe, the makers of Jev. Models are available on Hugging Face: Jeff-Qwen3.5-0.8B, Jeff-Qwen3.5-2B, and Jeff-Gemma4-E2B.