Jeff-0.8B: Fast Local Decision Models for Real-Time AI

GitHub - firelex/jeff: Fine-tunes of Qwen3.5 and Gemma 4 for zero-shot classification

There's something quietly subversive about a 0.8 billion parameter model that can make a decision in 30 milliseconds on consumer hardware. Most of us have internalized the idea that bigger models are better, that real capability requires real computational muscle. But here's a model that's smaller than a lot of smartphones' on-device speech recognizers, yet it's supposedly doing zero-shot classification with the same API-style interface that the much larger models use.

The numbers look almost too clean on paper. An 8B parameter version decides between 29 and 49 milliseconds per move on an M4 Max, while the public Doom-playing run from Jev clocked in at 212 milliseconds per call over its API. Those times weren't measured on the same hardware, so take the comparison with a grain of salt. But the broader point holds: something that fits in your laptop's memory is making meaningful predictions in real time.

What makes this interesting isn't just the speed or the size. It's the interface. You describe a situation, list your options in plain words, and the model returns a probability for each one. No fine-tuning, no training data that knows your specific categories. Whether you're routing support tickets, classifying user intents, moderating content, or just trying to get an AI to play Doom, you're talking to it like a person. That's the kind of flexibility that usually requires models an order of magnitude larger, and it's happening here in something that feels almost embarrassingly small.

What Jeff-0.8B Actually Does

Jeff-0.8B is a decision model, not a reasoning model. It takes a question with multiple options and returns a calibrated probability for each one in a single forward pass. No text generation, no parsing, no chain-of-thought. The output is just a vector of probabilities, one per choice. This is what makes it fast enough for real-time decisions on consumer hardware.

It handles three question types: choice (pick from A/B/C/D), noul (yes/no/uncertain), and score (rank options by preference). Each gets a probability distribution in one pass. On an RTX PRO 6000, it decides in about 28 ms. On an Apple M4 Max with MLX, it's 29–49 ms per move. For comparison, Jev's Doom run — the model it's based on — took 212 ms per call on a 32-thread CPU. That's a 4–7x speedup just from the architecture change.

The trade-off is that Jeff-0.8B doesn't generate explanations or intermediate reasoning. It's a pure classifier with calibrated confidence scores. You can run it locally with a few lines:

python -m mlx_lm.utils --model wfzyx/Jeff-0.8B-MLX \
  --prompt "Question: What is 2+2? A) 3 B) 4 C) 5" \
  --max-tokens 1

It was benchmarked on 599 questions across five public benchmarks, including JevBench's hard tier (105 items). The voice-navigation fine-tune is notable: held-out accuracy jumped from 31.7% to 95.8% after training on audio command pairs. That kind of gain suggests the base model had the capacity but needed the right task framing.

One thing that stands out: Von 1.2 reportedly had a better Doom score. Jeff-0.8B prioritizes speed over raw performance. Whether that's the right trade-off depends on your use case. For real-time applications, it probably is. For batch evaluation, maybe not.

Performance Benchmarks

On an M4 Max, Jeff makes a decision in 29–49 ms. On an RTX PRO 6000, it's 28 ms. Both measurements come from a single forward pass that outputs a calibrated probability distribution over the available options — no text generation, no parsing, just one pass through the model and a softmax.

That timing holds across 599 questions drawn from five public benchmarks, including JevBench's hard tier with 105 items and a set of 200 questions averaging about 200 input tokens each. For context, Jev's published Doom run took 212 ms per API call. This isn't just faster — it's the kind of speed that lets you run locally without thinking about latency.

The benchmark numbers tell a more nuanced story. Voice-navigation fine-tuning moved held-out accuracy from 31.7% to 95.8%, which is a solid improvement but also highlights how brittle the baseline was. The quote from von 1.2's Doom score — "had a better Doom score :D" — is a fair reminder that raw speed doesn't always win.

Here's how you'd time a single decision on your own hardware:

import time
import mlx.core as mx
from transformers import AutoTokenizer, MLXModel

model = MLXModel.from_pretrained("wfzyx/jeff-0.8b")
tokenizer = AutoTokenizer.from_pretrained("wfzyx/jeff-0.8b")

inputs = tokenizer("What is the capital of France?", return_tensors="mlx")
start = time.perf_counter()
logits = model(**inputs).logits
probs = mx.softmax(logits[:, -1], axis=-1)
mx.eval(probs)
elapsed = (time.perf_counter() - start) * 1000
print(f"Decision time: {elapsed:.1f} ms")

One thing that's genuinely confusing: the spec says "no parsing: ab," which I think refers to the answer format being a single letter (a, b, etc.) rather than a full text response. That detail matters because it's the difference between a classification head and a generation pipeline. The 28 ms on RTX PRO 6000 and 29–49 ms on M4 Max aren't just hardware differences — they reflect the architecture's efficiency at producing a single token's worth of information.

The CPU configuration uses 32 threads, but the real story is the gap between local inference and API calls. Jev's 212 ms per call includes network overhead, queueing, and whatever abstraction layer sits between the request and the model. Running Jeff locally cuts that down to under 50 ms, which is fast enough for interactive use on consumer hardware.

Practical Implementation

Integrating Jeff-style models into applications is straightforward in practice, even if the theoretical underpinnings are dense. These models produce a calibrated probability distribution over a fixed set of options in a single forward pass — no generated text, no parsing of output strings. That means you load the model, pass in your tokenized input, and read the logits directly. The output is a probability vector over your action space, which you can then map to whatever decision your application needs.

For real-time systems, latency matters. On an RTX PRO 6000, a single forward pass for Jeff-0.8B takes 28 ms. On an Apple M4 Max using MLX, it's 29–49 ms per move. That’s fast enough for interactive applications, but it depends on your batch size and how aggressively you optimize your inference pipeline. Running on a 32-thread CPU? You’re looking at roughly 200 ms per call, which matches the published Doom run that took 212 ms per decision step. Not ideal for high-frequency control, but perfectly usable for turn-based or semi-interactive agents.

The tradeoff becomes clear when you compare against API-based approaches. Cloud APIs add network round-trips — often 100–500 ms depending on region and load — on top of whatever compute time the remote endpoint needs. If your application can tolerate that, you avoid the hardware cost and maintenance burden. But if you need sub-50 ms decisions, local inference is the only option. You’re paying in hardware instead of in latency, and the math works out differently for everyone.

Benchmarking on 599 questions from five public benchmarks shows that Jeff-style models hold their own on structured decision tasks, especially when fine-tuned. One voice-navigation fine-tune moved held-out accuracy from 31.7% to 95.8% — a reminder that the base model is a starting point, not a finished product. The JevBench public hard tier (105 items) and a separate 200-question benchmark with ~200 input tokens each both point to the same conclusion: these models are good at picking from a known set of options, and they get better fast with the right training data.

Von 1.2 had a better Doom score :Dhttps://github.com/wfzyx/von — which is either a joke or a real comparison depending on how much you trust GitHub comments. Either way, it captures the spirit of this space: results vary, implementations differ, and the only way to know what works for your use case is to test it locally.

Can we get a price comparison? Edit: Running them for the masses. That question keeps coming up, and it’s the right one. An RTX PRO 6000 runs around $6,000. An M4 Max machine starts at $3,500. Cloud inference costs pennies per call but adds up if you’re making millions of decisions. The math changes fast once you factor in throughput.

The Tradeoffs

The system returns probabilities for each option rather than just picking a winner, which is genuinely refreshing after years of forced binary decisions. I can see this being useful for downstream systems that need to weigh confidence, or for humans who want to understand what the model was actually torn between. But it also means you're trusting the model's self-assessment of uncertainty, which has never been perfectly calibrated.

The noul question type — yes/no returned as a probability — feels like the most straightforward application. That's where confidence scores actually map cleanly to real-world decision making. If something is genuinely 90% likely to be true, that's actionable. The score question type is messier because it depends on the model describing its own scale, which introduces another layer of subjectivity.

I'm curious whether developers will actually use the full probability distribution or just take the top option like they do with traditional classifiers. My suspicion is most production systems will collapse this back to a single choice within weeks of integration, which would defeat a lot of the purpose. The real value here is in applications that can maintain and propagate uncertainty — something most software stacks aren't built to handle.

Conclusion

The numbers don't lie: 28 ms on an M4 Max, 49 ms on the same chip for the 8B variant, and a single forward pass that spits out calibrated probabilities for every option. That's fast enough to slot into real-time systems without architectural gymnastics. But I'm still not sure what to make of zero-shot decision-making that feels this frictionless.

Jeff-0.8B isn't trying to generate poetry or explain quantum physics. It's a decision engine with a probability distribution, and that limitation is its strength. You describe the situation, list your options in plain words, and it picks — no fine-tuning, no parsing, no generated text to clean up. For support queues, moderation labels, or game moves, that's genuinely useful.

The tradeoff is obvious: you're trading generative flexibility for speed and reliability. Whether that's a fair exchange depends entirely on your use case. For real-time inference on local hardware, it might be the difference between a prototype that works and a system that actually ships.