Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens
40% fewer tokens, same reasoning quality. Fireworks' Ember-1 post-trains Kimi K3 to compress output without cutting depth—a production win for teams paying by the token.

Why it matters
Fireworks released Ember-1, a post-trained variant of Kimi K3 that reduces token consumption by ~40% (from 49.3K to 29.9K output tokens per task) while maintaining reasoning quality. Now available as API-only Research Preview at Kimi K3 pricing. This matters for practitioners optimizing LLM costs in production without sacrificing capability.
The key facts
15 to knowEmber-1 is post-trained Kimi K3 variant
40% reduction in output tokens (49.3K → 29.9K per task in production A/B test)
Reasoning quality maintained at 'essentially unchanged score'
API-only Research Preview availability
Priced at Kimi K3 rates
Token compression via shorter reasoning traces, not reduced reasoning effort
Released by Fireworks AI
Production A/B test data provided
Fireworks AI released Ember-1, a post-trained version of Kimi K3
40% reduction in output tokens per task in production A/B test
Output tokens per task: 49.3K → 29.9K
Task score unchanged in A/B test
Available as API-only Research Preview
Post-training optimizes reasoning trace length, not reasoning effort
Published: September 28, 2026
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: Fireworks AI has released Ember-1, a post-trained Kimi K3 that learns to produce shorter reasoning traces instead of lowering reasoning effort. Fireworks reports about 40% fewer tokens, with output tokens per task falling from 49.3K to 29.9K in a production A/B test at an essentially unchanged…