September 28, 2026
Coding agents can use fewer tokens: Ember-1 cuts token use by 39%
According to Fireworks data from Sep 23, Ember-1 cut total tokens by 39% in a live A/B test on coding-task traffic. Its score was 0.753 versus Kimi K3's 0.751, while using 71.3% fewer reasoning tokens. Fireworks built Ember-1 on Kimi K3 and fine-tuned it to avoid repeating the same reasoning.

Cline
@cline
Ember-1, created by Fireworks' research team, is built on Kimi K3 and uses about 40% fewer tokens while achieving the same benchmark results. This was achieved through post-training K3: the model repeats the same reasoning less often. Reasoning models spend most of their output tokens, sometimes more than 90%, thinking before answering. In agent loops, this is costly: at every step, the model runs through the same thoughts again. Fireworks RL trained the model on real coding-agent task loops. This teaches the model to distinguish reasoning that changes the answer from going in circles. In a live A/B test on coding-task traffic, Ember-1 used 71% fewer reasoning tokens and 39% fewer tokens overall, while maintaining the same task success rate as K3.

· 2.5K views
Previously, on every new agent turn, the model would run through its reasoning again, including repeated steps. Fireworks trained Ember-1 on real coding-agent task loops to filter out thoughts that do not change the answer.
The model is already available. Ember-1 runs in Fireworks Serverless as `accounts/fireworks/models/ember-1`; create a key in the dashboard, then send requests through an OpenAI-compatible endpoint. It costs $3 per million input tokens and $15 per million output tokens, with a context window of 1.04 million tokens.
Ember-1 launched as a two-week Research Preview, after which Fireworks will decide whether to keep the model available permanently.
Source
