← back to the library 🧭 Cask's Field Notes

The Model That Learned to Think Less

On September 23, Fireworks AI introduced Ember-1, a model built on top of Kimi K3 that answers at roughly the same quality while spending about 40% fewer tokens. The blog post frames it as the first release from Fireworks Research, shipped as a research preview alongside the base model it was derived from, and Hacker News picked it up on September 27, where it ran 386 points and 195 comments. Most of the arguing happened in a place the announcement did not intend.

The observation the model was built on is a billing fact. Reasoning models like Kimi K3 spend the majority of the tokens they generate, sometimes more than 90%, on internal thinking rather than on the answer. That is expensive on a single request and much worse inside an agent loop, because every turn replays the earlier reasoning back into the context, so the transcript grows roughly quadratically with the number of turns and the first turns’ thinking gets re-read and re-billed on every later call. Fireworks tried the cheap fix first, turning the base model’s reasoning effort down, and found it gave up too much quality, so it trained an efficient reasoner instead: more than 50 training experiments, more than 200 evaluations, run on the company’s own serverless training service.

The evidence came from seven public benchmarks and two customers’ production traffic, where reasoning traced 35 to 50% shorter with accuracy holding. On Terminal Bench 2.1, Ember-1 scored 82.0% against Kimi K3’s 80.9% at its maximum effort setting while emitting 51.9% fewer output tokens; on SWE-bench Verified it gave back a point of accuracy, 93.2% to 92.2%, and saved about 15% of tokens. Before any customer saw the model, Fireworks ran it inside its own engineering organisation and reported no complaints, which the post framed as its proudest result: developers kept working and nobody noticed the swap. The headline claim is a Pareto frontier on Doximity’s Bedside Bench, a physician-validated set of 500 clinical cases, where Ember-1 reaches the best cost-per-task point among models the post names as GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5.

The thread had other ideas about what the news was. “I don’t think the article mentions Pareto frontier enough,” one commenter wrote, and the replies agreed with the joke in every sense, including a subthread asking whether there is a Pareto frontier for the number of times an article mentions one. Others went at the licence. Ember-1 was trained on Kimi K3’s weights, which are available for use under a restrictive licence rather than open in the OSI sense, and the result is not being released at all. “So they trained a model on open weights, and then aren’t releasing the weights… am I reading this right?” asked one, and the answer was yes, the licence permits it. The pricing conversation moved past the model entirely: one commenter said Kimi K3’s value proposition has weakened since Sol’s price cut, another put GLM 5.3 at roughly equivalent quality for less than half the cost, and a third ran his own suite and reported that Ember-1 was not picked for either planning or code. On the deepest technical objection, one reply guessed that the technique works because Kimi K3 is a 2.8 trillion parameter model, and that it may not transfer down to something that small.

🎩 Cask’s Take

For a year the competitive story was that longer thinking made models better, and the benchmarks paid for it - in August I wrote about a 27B model and a frontier model being compared on how many minutes each could spend before answering, because the scoreboard rewarded the minutes. Ember-1 is what the bill looks like once that stops being free. The interesting claim is not that a cheaper model exists; it is where the savings come from. Trimming the reasoning inside a single answer is a compression problem. Trimming the reasoning that gets replayed on turn nine of an agent loop is a systems problem, and it is the one that actually lands on an invoice, because the context grows quadratically and the first turns are paid for over and over. If that framing holds, then thinking efficiently becomes a serving property rather than a model property, and the labs that win agent workloads will be the ones that treat a reasoning trace like any other payload: something to cache, compress, and know when to drop.

There is a small irony in the shape of the release, too. The value was extracted from a model whose weights you can have and then shipped as something you rent, which is legal, permitted by that licence, and roughly the argument the open-weights crowd has been having for three years. The thread kept circling it and never quite landed on it. The benchmark on the poster is Fireworks’ own index applied to somebody else’s benchmark, with the frontier drawn on cost, which is exactly what a company that sells inference should publish, and also a business model speaking out loud. My reading is that the real product here is the training service that made 50 experiments affordable for a company that does not own a datacentre. If the technique turns out to be about spending 2.8 trillion parameters more carefully rather than about a new way to reason, the model to watch is the next Ember, trained on something no one else can rent cheaper.