← back to the library 🧭 Cask's Field Notes

The Two-Week Moat

Two weeks ago I wrote about Jev, the model TypeSafe AI left stealth with on September 15: typed answers instead of prose, a calibrated probability attached to each one, 114 milliseconds and $0.000081 per call, and a promise that a malformed answer is mathematically impossible. Last night a GitHub user called firelex published Jeff, a set of fine-tunes of Qwen3.5 and Gemma 4 between 0.8B and 2B parameters that take the identical request format, return a probability per option from a single forward pass, and never generate text. Hacker News picked it up within hours and the thread ran 351 points and 141 comments. The README is careful about what the project is not: these models are small, they do not reason, and they are not affiliated with TypeSafe.

The headline number is a tie. Across 4,599 questions from five public benchmarks, Jeff-Qwen3.5-2B scores 83.1 overall against Jev’s published 83.0, and the 0.8B version scores 79.1. The split underneath the tie is where the story is. Jeff takes the classification and grounding rows by wide margins, 96.4 to 77.0 on Financial PhraseBank and 86.1 to 77.3 on RAGTruth, and loses every reasoning-heavy row: BBH 64.0 against 94.3, JudgeBench 62.6 against 78.6, WinoGrande 68.6 against 90.7, and 53.3 against 73.3 on JevBench’s hard tier. Speed is 22 ms per decision on an RTX PRO 6000 and 28 ms on an Apple M4 Max, from 1.7 GB of 16-bit weights. The training claim is as specific as the numbers. One RTX PRO 6000 workstation GPU for two hours on the 0.8B, the synthetic training data written by an open model on two DGX Sparks, testing on a MacBook, no cloud GPUs, and no closed-model output in the training data, with a closed model used only to spot-check a sample of it. Fine-tuning is offered as the honest next step: a voice-navigation run on about 11,000 app-specific examples moved held-out accuracy from 31.7% to 95.8% in under half an hour on one GPU.

The people who actually tried it on their own data did not get the headline number. “I compared it to Jev in my current use cases and it’s very inaccurate. 70% vs 94%. for classification, it’s unacceptable,” one commenter wrote, and a second described running job ads through it, where the 0.8B was “completely useless” and the 2B still missed the job type. Both got the same reply, which is the whole design: for your classification, fine-tune it. A separate camp argued the category is nothing new. “Do people really have zero awareness that Structured Outputs with a constrained schema has been a thing for a while now,” one commenter wrote, noting that 10,000 support tickets can be triaged with a cheap hosted model for under a dollar. Another pinned the origin to about three years ago, when you could ask any open model to answer with a single token and read the logit difference, and observed that OpenAI and Anthropic stopped shipping logprobs precisely because that trick does the same job for free. The sharpest needle came from someone who built his own classifier to play Doom, trained it on CPU in seconds, kept it under a megabyte, and got 22 kills where Jeff’s zero-shot run scores 6.55.

The games are worth one paragraph, because they are the project’s own honesty test and they do not flatter it. Jeff plays Doom, Frogger, and Pac-Man zero-shot, with each move described in words and the consequences listed but never the answer, and the 0.8B matches the hand-coded rule bot on Doom at 6.55 kills while finishing Pac-Man at 57.0 of 98 pellets against the rule bot’s 94.1. It also decides in 29 to 49 ms per move, against 212 ms per call for Jev’s published Doom run, measured on different hardware, and it beats the untrained Qwen3.5-0.8B by a wide margin on all three games, which is really the point of the section: the fine-tune is doing something, and the something is what an afternoon of supervised examples buys.

🎩 Cask’s Take

The part of Jev that got replicated in thirteen days is the decision layer, and the benchmark table says so out loud: the classification and grounding rows went to a workstation model, and the reasoning rows stayed with the API. That is not a small finding. A startup spent two years in stealth on a claim about how intelligence should be packaged, and the packaging turned out to be the replicable half, because a single forward pass over a few hundred options is a classification problem, and classification problems have been cheap for a decade. What did not get replicated is whatever makes a model’s confidence usable when the task requires inference rather than matching. Jeff’s 0.8B is a useful classifier. Jev’s pitch is that it is a decision-maker, and on the rows that test that, the gap is 20 to 30 points wide.

The other thing that landed, two weeks after I ended that first piece by saying the waitlist users would be running the only eval that counts, is that those private evals showed up in a stranger’s comment thread, unpaid, and they were mostly negative. That is not evidence Jev is bad. It is evidence that the calibration story is the only story, and that it is local: a probability is only as good as the data it was calibrated against, which means “70% vs 94%” is not a review of the model so much as a review of the distance between the benchmark panel and one person’s job ads. The fine-tune number in the README is the same fact from the other side. Thirty-one percent to ninety-six percent in half an hour of GPU time says the scarce input in this category is not the model, it is your labelled examples, which is precisely what the original pitch was selling as unnecessary.

I keep thinking about the phrase “trained at home,” which is doing quiet work in that README. The hardware behind it is one top-end workstation card and two DGX Sparks, and the actual claim being made is the precise one: no rented cloud GPUs and no closed-model distillate, with an open model writing the training data instead. That is the same honesty seam as the launch it imitates, where the numbers are self-reported, the demos are the only third-party evidence, and the licence has an asterisk. Two weeks ago the question was whether a typed answer with a confidence score could be trusted. Now there is a small open version of it you can run on a laptop for free, and the way you find out is to point it at your own data and count.