He Trained a 0.8B AI at Home. It Decides in 22 ms. Nobody Saw This Coming.

Bottom line: Firelex released Jeff, three small decision models fine-tuned from Qwen3.5 and Gemma 4. They return calibrated probabilities over user-defined options in a single forward pass.

The 0.8B variant decides in about 22 ms on an RTX PRO 6000 and 28 ms on an Apple M4 Max, and scores 79.1% across five public benchmarks. The 2B model scores 83.1%, just ahead of Jev at 83.0%.

For routing, moderation, and intent classification, you may no longer need an API call per decision.

The Most Expensive Line in My Stack Was a Classifier

I've shipped a lot of systems where the "smart" part was a frontier model and the actual work was a `switch` statement. Is this ticket billing or technical?

Should this request go to the cheap model or the expensive one? Is this message safe to pass through?

I paid frontier-model prices and waited hundreds of milliseconds for answers a much smaller model could have given. I knew that, and I did it anyway, because nothing small was good enough to trust.

That is the trade Jeff attacks.

Jeff hit Hacker News on September 29, 2026, and it's sitting near the top of the front page with a few hundred points and a long comment thread.

I've read the release material, not run the models myself. So this piece is an analysis of the published numbers, and I'll say where they stop being enough.

What Jeff Actually Is

Jeff is a trio of small models from Firelex, fine-tuned from Qwen3.5 and Gemma 4. You hand each one a situation and a list of options you define.

It returns a calibrated probability for each option, in one forward pass.

There's no chain-of-thought and no token-by-token generation. Nothing has to be parsed out of prose. The model reads the input and scores your options directly.

That design is why it's fast. A generative model that "thinks out loud" before answering pays for every token. A model that scores options pays for one pass.

The release is also described as Jev-compatible. I take that to mean it fits into the workflow of Jev, the existing decision-model baseline it's compared against below.

I haven't verified the exact interface, so check the repo before you plan a migration.

The Numbers, Read Carefully

Here's what has been reported:

Article illustration

The 2B result is narrow. 83.1% against 83.0% is a tie for practical purposes, and I wouldn't build a purchasing decision on a tenth of a point.

The real story is that a home-trained model matches the reference on these benchmarks while being small enough to run on a laptop.

The 0.8B's 79.1% is about four points lower. Whether that gap matters depends on your problem, and I'll get to that.

Why the Training Story Matters More Than the Speed

Two hours on one workstation GPU is not a research-lab budget. It's a long lunch and an electricity bill.

For years the assumption was that useful decision-making needed either a giant model or a giant training run. Jeff's numbers suggest a third route.

Take a capable small base model, narrow the task hard, and fine-tune it on your own hardware. Narrowing the task is what makes the economics work.

A model that only has to rank options doesn't need to be good at poetry or trivia.

This also changes who can build these things. If a 0.8B model takes two hours to train, then a team with a domain-specific decision problem can train its own.

You don't have to wait for a vendor to guess your categories.

Where This Fits in a Real System

Think about every place your product currently makes a cheap judgment call:

Each of these is a decision over a fixed set of options. That's exactly Jeff's shape.

At about 22 to 28 ms for the 0.8B model, depending on hardware, these calls can sit inline on the request path instead of being an async afterthought.

Calibration deserves more attention than it gets.

A probability you can trust lets you build thresholds: act automatically above 0.9, ask a human between 0.5 and 0.9, escalate to a bigger model below that.

An uncalibrated model that says "90%" about everything gives you nothing to threshold on. If Jeff's probabilities hold up on your data, that's the feature that makes it usable in production.

And as I wrote about when agents kept agreeing with everything, a model that sounds confident is not the same as a model that's right. Calibration is how you measure the difference.

The Reality Check

I don't want to oversell this, so here's where I'd slow down.

Public benchmarks aren't your data. Five benchmarks tell you the model is competent in general.

They don't tell you how it handles your weird category boundaries, your jargon, or your adversarial users. Test on a few hundred of your own labeled examples before you trust it.

The 0.8B trade-off is real. Four points of accuracy is cheap on a low-stakes routing decision and expensive on a moderation call where a miss has consequences. The 2B model exists for that reason.

Pick the size against the cost of being wrong, not the speed chart.

Single-pass scoring has limits. It's great for choosing among options you defined. It can't do open-ended reasoning, and it can't tell you "none of these fit" unless you give it that option.

Some decisions need the long thinking a bigger model does.

I haven't benchmarked it myself. Everything above comes from the published results.

Independent reproductions will matter more than the launch post, and the Hacker News thread is a good place to watch for them.

What I'd Do This Week

If you run any system with an LLM call whose only job is choosing among a few options, try this:

Article illustration

1. Inventory your decision calls. Grep for every place a model returns a label, a route, or a yes/no. Count them and note the per-call latency and cost.

2. Build a labeled test set. Pull 300 to 500 real examples with the answers you'd have wanted. This is the step people skip, and it's the only one that matters.

3. Run the 0.8B and 2B side by side. Compare against your current API model on accuracy, and check whether the probabilities are calibrated by bucketing them and comparing to actual hit rates.

4. Design the fallback ladder. Use Jeff for confident cases and escalate low-confidence ones to your larger model. You keep the quality and drop most of the bill.

5. Consider fine-tuning your own. If two hours on one GPU is the real cost, your domain labels may beat a general model.

The pattern I'd bet on is not "small model replaces big model." It's a small model handling the 80% of easy decisions instantly, with the big model kept for the hard remainder.

The Bigger Point

The most interesting thing about Jeff isn't a single benchmark score.

It's that decision-making, the part of AI systems we assumed was expensive, is compressing down to something a workstation can produce in an afternoon. That shifts where the leverage sits.

The advantage goes to the people with good labeled data about their own problem, not the people with the biggest API budget.

Which of your API calls do you think is really just a classifier in disguise, and would you trust a 0.8B model with it?

Sources: Jeff on Hacker News, AI Weekly: Firelex ships Jeff, AI/TLDR release note

Story Sources

Hacker Newsgithub.com