Interestana
Home/News/Fireworks AI Releases Ember-1 With 40% Fewer Tokens
MarkTechPost••3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Fireworks AI Releases Ember-1 With 40% Fewer Tokens

Fireworks AI has released Ember-1, a specialized model developed by Fireworks Research through post-training Moonshot AI's open-weight Kimi K3. This new model is designed to produce shorter reasoning traces while preserving task accuracy, a feat achieved by fundamentally altering the model's reasoning process rather than simply adjusting inference settings. According to a release post from Fireworks, Ember-1 delivers the quality of Kimi K3 with approximately 40% fewer tokens. Currently, Ember-1 is available as a Research Preview exclusively through the Fireworks serverless API. Fireworks has not released the model's weights, training code, or specific training algorithms, making self-hosting impossible at this time.

The development of Ember-1 addresses a significant challenge with current reasoning models, where they can expend over 90% of generated tokens on internal reasoning processes. This inefficiency becomes particularly costly in multi-turn agentic workloads, where prior reasoning is replayed to the model in each subsequent turn. The context window grows roughly quadratically with the number of turns, leading to long reasoning traces from early turns being re-read and re-billed on every later call. Fireworks stated that customers desired Kimi K3's coding capabilities at a reduced cost, but lowering K3's reasoning effort setting resulted in an unacceptable loss of quality. Consequently, the Fireworks team focused on training the model to reason more efficiently.

Fireworks Research constructed Ember-1 by retaining useful self-reflection behaviors, such as revisiting assumptions or responding to feedback, while eliminating redundant reasoning and unproductive loops. The training dataset encompasses a broad range of tasks, including mathematics, coding, instruction following, conversation, search, tool use, and software engineering. This dataset includes both standalone problems and extended multi-step interactions. The training process utilized task and environment feedback to guide on-policy planning and learning. The Fireworks team conducted over 50 training experiments and more than 200 evaluations. They also developed new, unpublished training algorithms. All training operations were performed on Fireworks' own infrastructure.

This innovation by Fireworks AI aims to make advanced reasoning models more cost-effective for applications that require extensive sequential processing or agentic behavior. By reducing the token overhead associated with internal thought processes, Ember-1 could enable more complex and longer-running AI agent tasks without a proportional increase in computational cost. The focus on maintaining accuracy while reducing token count is a critical step towards more efficient and scalable AI deployments, particularly in competitive cloud API markets where cost per token is a key metric for adoption. The model's availability as a research preview suggests that Fireworks AI is seeking feedback and further refinement before a broader commercial release.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next