Fireworks says Ember-1 keeps the answer and cuts the token bill
The Fireworks Research model is built on Kimi K3 and trained to reduce unnecessary reasoning while keeping coding and agent quality intact.
What happened
Fireworks Research introduced Ember-1, a specialized model built on Kimi K3 that the company says delivers comparable quality with about 40% fewer tokens.
The company says it trained Ember-1 after more than 50 training experiments and over 200 evaluations, using Fireworks' own data rather than customer data.
Why developers care
Reasoning traces can dominate the cost of coding and agent workloads, especially when multi-turn systems keep replaying previous reasoning back into context. Fireworks argues that K3's excess reasoning could be shortened without removing the useful self-checking behavior.
In Fireworks' live A/B tests with two customers, the company says Ember-1 used about 35% fewer tokens per task at comparable quality. It is launching as a Research Preview on Fireworks Serverless, with training support also planned.
Quick takeaways
- Fireworks positions Ember-1 as a token-efficient Kimi K3-based model for coding and agentic workloads.
- The company reports reasoning-token reductions of 35-50% in benchmark and production settings.
- Ember-1 is available as a Research Preview, with permanence tied to community demand.