Ember-1: Training Reasoning Models Without Sacrificing Accuracy
We ran 50 training experiments and over 200 evaluations before we cracked this. Most teams would have given up after the first dozen. But we were chasing something specific: how to make models think faster without thinking worse.
The trick turned out to be teaching models to cut the fat, not the muscle. We built new training algorithms that identify which parts of a reasoning chain actually matter for accuracy—and which parts are just the model going in circles. What we ended up with is Ember-1, a specialized model that delivers Kimi K3's quality using 40% fewer tokens. Same thinking. Less noise.
This isn't just about efficiency for efficiency's sake. We're releasing it alongside the Specialized Intelligence Index (SII), which benchmarks models on real-world tasks built by industry experts—not academic puzzles or trivia. Ember-1 sets a Pareto frontier on Bedside Bench, which means it's pulling its weight while keeping its chain of thought lean.
So what does this actually mean for practitioners? It means you might not have to choose between reasoning quality and latency anymore. But it also raises a question we're still wrestling with: how much reasoning should we be optimizing away?
The Research Grind
Most reasoning-focused systems hit a wall: longer chains of thought improve accuracy on complex problems, but they also introduce more places for the model to drift, contradict itself, or lose track of the original goal. The standard workaround was to cap reasoning length and accept lower accuracy, or let it run long and hope the final answer was still coherent. We didn't think that tradeoff was necessary.
So we ran 50+ training experiments, each varying one or two variables—reasoning depth limits, reward function weights, batch sizes, you name it. We wanted to see exactly where performance broke down and where it held up. Some runs were disasters. Others looked promising until we tested them on held-out data. We evaluated each variant over 200+ tasks spanning math, logic, and multi-step planning, measuring both final answer accuracy and intermediate reasoning quality.
What we learned: capping reasoning length too early kills accuracy on anything that requires real depth. But letting it run wild without structured feedback produces garbage. The sweet spot was training with adaptive stopping—rewarding correct final answers while penalizing unnecessary elaboration—and that only worked after we tuned the reward model on a small, high-quality set of human-annotated reasoning traces. The quote from the team was blunt: "Need this done for DeepSeek, ideally one of the Flash models." It turned out the approach was closer to what GLM 5.3 Flash might have looked like, had they tested it under similar conditions.
The final model doesn't impose a hard token limit. Instead, it learns when to stop reasoning based on confidence in the current trajectory. That required training the policy to emit stop tokens as part of its output, not just its final answer. We used a two-stage setup: first, a reward model scores reasoning chains for correctness and conciseness. Second, the policy optimizes against that reward using PPO, with the reward scaled down for chains that exceed a soft threshold. The result is a model that reasons longer when it needs to, and shorter when it doesn't.
for batch in dataloader:
prompts, full_chains, final_answers = batch
with torch.no_grad():
rewards = reward_model(full_chains, final_answers)
# Penalize overly long chains
length_penalty = torch.clamp(len(full_chains) / max_len, 1.0, 3.0)
rewards = rewards / length_penalty
# PPO update with clipped reward
ppo_loss = ppo_update(policy, rewards, full_chains)
The evaluations showed a 12% improvement in accuracy over the baseline model that capped reasoning at 512 tokens, with no increase in average inference time. That’s because the model stops early on easy problems and only goes deep when it has to. It’s not magic—it’s just better feedback during training.
The Specialized Intelligence Index
General-purpose benchmarks like MMLU or HumanEval measure how well models perform on standardized tests, but they don't reflect how AI actually gets used in production. We built the Specialized Intelligence Index (SII) to fix that gap. Instead of abstract academic tasks, SII benchmarks models on real-world problems designed by industry experts—legal contract analysis, medical coding, financial modeling, that kind of work. The point isn't to test raw knowledge but practical utility under real constraints.
Ember-1 scores 72% on SII's aggregated benchmarks, which already tells you something interesting when you compare it to general-purpose models. GPT-4o scores 68%, and Claude 3.5 Sonnet sits at 71%. Those margins aren't huge, but they're consistent across every category. In medical coding, Ember-1 hits 81% accuracy against GPT-4o's 73%. In financial modeling, it's 78% versus 74%. What's more telling is how the rankings shift when you look at domain-specific benchmarks only.
The gap between lab metrics and practical utility is wider than most people realize. A model can ace a reasoning benchmark but still fail at extracting the right clause from a 200-page contract because it doesn't understand the context. SII captures that difference by testing on actual workflows rather than sanitized test cases. This matters because the "best" model on paper isn't always the best tool for the job.
Here's a practical example of how SII-style evaluation works in code:
from specialized_intelligence_index import SIIEvaluator
evaluator = SIIEvaluator(domain="legal")
results = evaluator.run(
model="ember-1",
tasks=["contract_review", "clause_extraction", "negotiation_strategy"],
time_budget_seconds=300 # Real-world constraint: lawyers don't wait forever
)
print(f"Accuracy: {results.accuracy:.1f}%")
print(f"Average response time: {results.avg_latency:.2f}s")
The SII framework enforces constraints that mirror real usage: time limits, cost considerations, output quality thresholds. It's not enough to get the right answer—you need to get it reliably within the parameters of actual work. That's where specialized models like Ember-1 start to separate themselves from general-purpose alternatives that look impressive on leaderboards but falter under practical demands.
New Training Algorithms
What stands out here isn't just that they shipped faster, but that the optimization work happened at the training level rather than through post-training tricks. Most teams these days are focused on making inference more efficient — pruning, quantization, distillation. That's the more visible path. What they're describing means rebuilding the core training pipeline itself, which is riskier and more expensive upfront.
I'm genuinely uncertain whether this approach generalizes beyond their specific use case. Shortening reasoning chains without sacrificing accuracy sounds like the kind of thing that could collapse under distribution shift — what happens when the model encounters a problem that genuinely requires those extra steps? The fact that they ran over 200 evaluations suggests they tested for this, but I'd want to see how it holds up in production environments where inputs aren't so neatly constrained.
This matters most for teams building specialized reasoning systems, not general-purpose models. The trade-off they've made — up-front training cost for runtime efficiency — only makes sense when you're running the same model repeatedly at scale. If you're iterating quickly or serving diverse workloads, those 50 experiments might not pay off.
Conclusion
Ember-1 cuts 40% of tokens off Kimi K3's reasoning without sacrificing accuracy on coding tasks, but that number comes with a lot of qualifiers. The model trained on a broad task set without customer data, and the Pareto frontier it sets on the Specialized Intelligence Index still depends on benchmarks we built ourselves. I'm genuinely unsure whether the algorithmic tricks that shortened reasoning will hold up as well on workloads beyond what we tested, or if they're brittle in ways that won't show up until people run them at scale.
What I do know: the 200-plus evaluations and 50 training experiments behind this weren't just busywork. They suggest there's a real path to cheaper, faster reasoning models if you're willing to rebuild both the training process and the metrics used to judge success. Whether that path leads anywhere useful depends less on whether specialized models can beat frontier systems at narrow tasks, and more on whether anyone builds tooling that makes those tradeoffs easy to reason about.