RightOne.ai

Loading your page.

RightOne.ai blog

RightOne.ai engineering notes

What RightOne.ai is optimizing: inside AutoRouter v1 orchestration

We are optimizing for the right answer at the right effort level under real product constraints: quality, latency, provider policy, cost, context, and user trust.

Published . Updated . 7 min read

Key takeaways

  • AutoRouter v1 treats each chat turn as a routing decision, not just a provider API call.
  • The orchestrator combines prompt shape, session context, model metadata, account policy, and provider behavior before dispatching.
  • The goal is not always the cheapest model or the strongest model. It is the route that fits the job safely and efficiently.

The RightOne.ai team is building ChatOS around a simple belief: users should not have to know which AI model is best for every turn. The model landscape changes too fast, and the best route depends on the current prompt, the conversation, the account policy, and provider availability.

What AutoRouter v1 evaluates

  • Prompt shape: coding, math, summarization, research, casual chat, planning, or mixed intent.
  • Difficulty and effort: whether the turn needs none/low effort, medium effort, high effort, or max/xhigh-style exploration.
  • Context needs: how much of the conversation should be carried into the next provider call.
  • Model metadata: context window, pricing, provider, strengths, limitations, and observed routing quality.
  • Policy: account state, privacy constraints, fallback rules, and provider eligibility.

Orchestration is more than picking a model

A production chat turn has several stages. The backend owns session identity and chat ownership. The router classifies the turn. Policy narrows eligible providers. The provider gateway prepares the request, clamps output limits, streams deltas, records safe public events, and reconciles usage where provider data is available.

This is why we avoid fixed per-message thinking. A short rewrite and a deep architecture review should not reserve or reconcile the same amount of compute. Provider usage and model pricing matter.

Why effort-aware routing matters

Effort-aware routing lets AutoRouter v1 avoid two common failures: overkill and underkill. Overkill sends easy tasks to expensive high-effort routes. Underkill sends hard tasks to cheap routes that fail silently or produce shallow answers. The route should move with the task.

What we optimize for

  • Quality: choose a model likely to answer the task well.
  • Efficiency: avoid premium routes when lighter routes are enough.
  • Latency: keep simple turns fast and reserve high/max effort for hard work.
  • Privacy and policy: do not expose private traces or rely on browser IDs for ownership.
  • Continuity: preserve account-backed history and use context only when it improves the turn.

Where this is going

RightOne.ai is laying the foundation for account-owned chat history, publishable conversations, profiles, collaboration, branches, forks, and contribution-style workflows. But the first priority remains the core loop: a user asks, ChatOS safely starts a turn, AutoRouter v1 chooses the route, and the answer streams back with clear public status events.