September 2026. Every number here is from the benchmarks, and bash experiments/bench.sh --no-record reruns them without an API key. If you classify text with Jev, every answer is a network call to one vendor and comes back in about 300 ms, at any load. For an agent loop or game tick, 300 ms per step is the whole budget. Jevstiller sits in front of that call, learns a small local model from Jev’s own answers, and lets it answer what it is sure about in about 15 ms on a CPU.
The interesting part is not the small model. It is the contract: Set one number, say 98%. Jevstiller returns the label Jev would have returned on at least that share of requests. This post is about what it takes to make that sentence true, why the obvious way of picking a confidence threshold does not make it true, and what it costs.
The usual recipe for a cascade like this: hold out some data, sweep a confidence threshold, keep the loosest one whose measured disagreement is within budget, ship it. Then it breaks the budget about half the time. On five public tasks, twenty random train/calibration/test splits each, the point-estimate rule exceeded the 2% budget on 6 to 12 of 20 splits per task, by up to a full percentage point. For example, on Banking77, the point estimate rule had mean disagreement 1.90% and worst 2.70%, breaking budget 9/20 splits, while the bound rule had mean 1.15% and worst 1.55%, breaking budget 0/20.