Effort-aware AI
Effort level in AI routing: from none and low to xhigh, max, and ultracode
Effort level is the provider control that changes how much reasoning budget a model spends. Labels differ by lab and product: default, none, low, medium, high, xhigh/extra-high, max, and Claude Code’s ultracode are not one universal scale, but they describe the same tradeoff: speed versus deeper work.
Published . Updated . 6 min read
Key takeaways
- Effort should match the task: quick lookup, careful explanation, coding/debugging, and proof-style reasoning need different provider settings.
- More reasoning can increase latency and token usage without improving simple answers.
- AutoRouter v1 is designed to estimate when to escalate effort and when to keep the route lean.
Most people think model selection means choosing the biggest model they can afford. In practice, the effort setting can matter just as much as the model name. A tiny factual question should not burn high or max reasoning. A multi-file debugging task should not be routed like a casual one-sentence question.
The familiar provider effort variants
The exact names are provider-specific, but the common ladder looks like this: default, none, low, medium, high, max, and sometimes xhigh or extra-high. Claude Code also introduced a product-specific ultracode setting around Opus 4.8 dynamic workflows; Anthropic describes ultracode as setting effort to xhigh while letting Claude decide when to use a workflow.
Default
Default means the provider or product chooses the baseline. It is not necessarily medium. For example, Anthropic says Opus 4.8 defaults to high effort because they judge it to be the best balance of quality and user experience.
None or low
None or low is for direct answers, short transformations, formatting, everyday explanations, and tasks where latency matters more than elaborate reasoning. The answer should be concise and cheap enough to use frequently.
Medium
Medium is for mixed tasks: moderate context, some reasoning, some summarization, and a few constraints. It is often the safest general-purpose setting when the task is not trivial but does not justify high-effort exploration.
High
High is for harder reasoning, coding, debugging, architecture decisions, ambiguous requirements, legal-style comparisons, and multi-step planning. It can improve answer quality but usually spends more tokens and time.
Max, xhigh, extra-high, and ultracode
Max or xhigh/extra-high is for work where mistakes are expensive or the task is long-running. Claude Code’s ultracode goes beyond a simple answer setting: it combines xhigh effort with dynamic workflows, where Claude can fan work out across subagents and verify results. That is powerful, but it is not a default for everyday chat.
Why blindly maximizing effort backfires
- Latency grows: the user waits for reasoning that may not change the final answer.
- Token usage grows: visible or hidden reasoning consumes budget.
- Style can worsen: deep models may over-explain, hedge, or repeat intermediate checks.
- Opportunity cost rises: using a frontier route for easy work leaves less budget for hard work.
How RightOne.ai thinks about effort
RightOne.ai reads the shape of the prompt: subject, ambiguity, task type, expected answer length, context needs, and recent session state. It then translates the user’s need into provider-specific routing language: maybe none/low for a quick rewrite, medium for a mixed explanation, high for debugging, or max/xhigh-style effort for high-stakes multi-step work.
This is why a chat platform needs orchestration. A single conversation can contain a quick definition, a code review, a long research question, and a precise rewrite. Treating those turns equally is wasteful.