Reasoning models
What is MoE, and why do reasoning models repeat “wait” and “actually”?
MoE is an efficiency architecture. Visible long reasoning is a post-training behavior. The two often appear together in modern open models, but they are not the same thing.
Published . Updated . 8 min read
Key takeaways
- Mixture-of-Experts activates only part of a model for each token, which can improve capacity-per-dollar when serving is engineered well.
- Long visible reasoning traces often come from reinforcement learning and distillation recipes that reward step-by-step self-correction.
- Closed products usually hide or compress scratch work, so they may appear more direct even when they internally reason through multiple steps.
MoE stands for Mixture-of-Experts. Instead of activating the entire model for every token, an MoE model contains multiple expert blocks and a router that selects a small subset of those experts per token. This can give the model a large total parameter count while keeping active computation lower than a dense model of the same size.
That sounds similar to RightOne.ai routing, but it happens at a different layer. MoE routing happens inside a model during inference. RightOne.ai routing happens before the provider call: it chooses which external model route should handle the user’s turn.
MoE in one paragraph
Imagine a company with many specialist teams. A dense model asks the whole company to review every word. An MoE model asks a dispatcher to send each word to a few relevant teams. If the dispatcher is good and the infrastructure is efficient, the company can be large without making every task expensive.
Why DeepSeek-R1-style models can think out loud
Open reasoning models such as DeepSeek-R1 popularized long visible reasoning behavior. The important point: the repeated phrases are usually not because the model is “confused” in a human sense. They are artifacts of training recipes that reward self-checking, revision, and step-by-step exploration.
- “Wait” often marks a correction point: the model is changing direction after spotting a contradiction.
- “Actually” often marks a re-evaluation: the model is revising an earlier assumption.
- Long chains can appear when a reasoning model has been rewarded for showing its work, even when a shorter answer would be enough.
This does not mean users should receive every internal scratch token. In many production systems, reasoning is hidden, summarized, or transformed into a concise answer. Visible reasoning can be useful for debugging and learning, but it can also waste tokens, increase latency, and overwhelm the reader.
Why closed-source models often answer more directly
Closed-source model products usually separate internal problem solving from the final user-facing response. Providers can train the assistant to be concise, hide private scratchpads, and format the output as an answer instead of a transcript of exploration. They also run large preference-tuning and safety-tuning loops that penalize unnecessary repetition.
That product layer matters. A closed model may internally evaluate several paths, call tools, or consult retrieval. The user sees a polished final answer. An open reasoning checkpoint served with minimal response shaping may show much more of the intermediate exploration.
Where MoE helps and where it does not
- MoE can improve serving economics for large-capacity models, but it requires careful routing, batching, memory, and hardware utilization.
- MoE does not automatically make a model truthful, concise, or safe. Post-training and product policy still matter.
- MoE does not replace external orchestration. You may still need to choose between a small dense model, a large MoE model, a frontier closed model, or a retrieval-heavy route.