What automatic routing saves and what it costs in predictability, which of eight products document it, and the settings that keep a router honest
Aug 18, 2026 · 13 min read
A team should turn on automatic routing when a real share of its prompts are extraction, rewriting, classification or simple drafting, and should keep a named model for regulated or high-stakes work. The setup that works is a governed workspace or gateway that logs the model it chose, allows an override, limits the eligible models and is measured on your own work. The eight workspaces compared on the same criteria below are Playgram, WorkLLM, nexos.ai, Langdock, TeamAI, Aymo, Magai and TypingMind.
The saving is real and it is workload-dependent. Model prices sit far apart, and TeamAI's own published rate card lists GPT-5 Nano at $0.05 per million input tokens against $21 for GPT-5.2 Pro35. A router earns its place by moving routine traffic down that gap, and it earns nothing when almost every prompt already needs a frontier model.
This guide sets out how the routers on the market decide, what the loss of predictability costs, and which of the eight products document routing at all. It ends with a pilot that measures quality on your own prompts rather than on a leaderboard.
Routine extraction and hard analysis arrive in the same day and go to the same model.
Credit or API spending has become material and you need to restrict without blocking.
Some outputs need a named model and a person checking them, whatever a router would pick.
One model handles nearly everything, or usage is too light for the economics to matter.
Every prompt goes to the strongest model or every person has to decide for themselves, and both of those have a price.
Separate subscriptions duplicate model access and keep charging for seats nobody uses, and application teams usually set one capable model as the default because that is operationally simple. The price gap that creates is wide. TeamAI's published rate card lists GPT-5 Nano at $0.05 per million input tokens and $0.40 per million output, against $21 and $168 for GPT-5.2 Pro35. On the subscription side, four single-vendor team plans came to about $101 per person a month at July 2026 list prices. That example total takes ChatGPT Business at $25, Claude Team at $25, Gemini Business at $21 and Grok Business at $301, 2, 3, 4. Read that as an illustration rather than a rate, because a cheaper mix is easy to assemble.
Without routing, people switch tabs, copy prompts, upload the same file again and test several models by hand. A comparison mode helps during an evaluation and it multiplies the calls, because every selected model answers the same prompt in full. Routing makes one choice before generation instead, which is a different purchase from comparing after it.
Moving a task to another product breaks the thread, so the earlier messages, uploaded files, project instructions and decisions have to be copied or summarised again. Inside one workspace this varies by product rather than being a category guarantee. Langdock lets a person switch away from the automatically chosen model during the conversation, and Aymo says a model can be changed without starting a new thread. Magai says chat history and uploads stay available across a change36, 22, 24. Routing adds a second requirement on top, which is that every eligible model receives the same retrieved project context, or the comparison between them is not fair.
An unmanaged router hands administrators a new visibility problem rather than solving one. They need to know which model answered, why it was chosen, what the call cost, whether a fallback happened and who started it, and whether an expensive model can be blocked before the request runs. A monthly aggregate saving does not answer any of those. nexos.ai documents per-request logs with use and cost by model, user, team and project, and budgets and hard caps set before spending occurs12.
Five checks on any product that offers to choose the model for you. The first two decide whether the router is governable and the rest decide whether it is useful.
The router needs genuinely different options on cost, speed and capability, and an admin needs to exclude models by region, security, capability or price. Look for a stated objective too, since Microsoft exposes cost, balanced and quality modes rather than leaving you to guess what it optimises.
Check each product for web search, deep research, document analysis, spreadsheet work, image generation, video generation, code explanation, chats that leave no history, and MCP connections. A cheaper model may not support the tool call, so the router has to filter those models out.
The chosen model should get the thread, the attachments and the project instructions, and each eligible model should receive the same retrieved context. Otherwise a cheaper model looks weaker only because it was given less, and the routing test measures the wrong thing.
A person should see which model answered and be able to rerun on a named stronger one. An admin needs usage by person, model and period, budgets or hard caps that act before the request, scoped permissions for sensitive projects, and a new starter who inherits approved models and instructions.
Routing changes what the models consume rather than what the seats cost, so a per-seat plan hides the benefit. Look for pooled credits or metered use where a cheaper route reaches the invoice, and check what a light user costs, since a fixed seat price undoes the saving on people who barely log in.
The multi-model workspaces a team is most likely to weigh up, on the same criteria and to one standard. Where a vendor does not document something, the cell says so.
This table compares multi-model team workspaces with each other. The single-vendor plans a workspace replaces are priced further down, under 'Priced per seat', and are not rows here. Infrastructure routers such as the Microsoft and Amazon services discussed on this page are not rows either, because they route application traffic rather than a team's chat. Pricing is the lowest-priced paid plan that covers five users, at the monthly rate. Each cell cites the page that documents that cell rather than one pricing page per row. Figures checked August 2026, and cells marked 'Manual test required' could not be confirmed from public documentation.
The same products on the criteria that decide daily use: what each one does besides chat, what it connects to, what an admin can see and cap, and where your data goes.
These criteria decide daily use more than the model list does, and vendors document them very unevenly. 'Not publicly documented' means the official sources checked did not state it, and it does not mean the feature is absent, so read those cells as questions to put to the vendor. Checked August 2026.
The published per-seat price of each major single-vendor team plan, billed monthly. None of them routes across another vendor's models, because each one only ships its own.
Buying all four for one person came to about $101 a month at July 2026 list prices. Read that as one example stack rather than a going rate, because a team can assemble a cheaper mix and the four plans do not buy the same amount of use. Routing changes the other half of the bill, which is what the models consume rather than what the seats cost, so compare both. The estimator further down runs that comparison on your own numbers.
Routing changes what the models consume, and it leaves the seat count alone, so two of these move and two of them do not.
Six places to put the choice of model, from a workspace that decides for you to a rule your own engineers write. The workspace options come first.
The product picks a model per prompt or per conversation and lets a person override it. nexos.ai documents Auto Select by quality and specialisation, and Langdock evaluates the first message and keeps that model for the rest of the conversation36, 37.
Best for: Teams with a mix of routine and hard work in one queue.
Strengths
Trade-offs
The same setup with the router given a short approved model list and one fixed model for regulated work. Microsoft exposes cost, balanced and quality objectives and supports custom model subsets, and nexos.ai documents model exclusions through guardrails32, 12.
Best for: Teams with regulated work beside ordinary drafting.
Strengths
Trade-offs
Standardise on a single vendor and use whatever model picker it ships. Behaviour is predictable and procurement is simple, and there is no routing across providers because there is only one provider.
Best for: Teams whose work fits one model family.
Strengths
Trade-offs
Buy the plans the work needs and let people choose by hand. As an example, all four plans priced above came to about $101 per person a month at July 2026 list prices, and each of them only routes inside its own family1, 2, 3, 4.
Best for: Teams that need each vendor's own tools daily.
Strengths
Trade-offs
A gateway routes the traffic your applications generate rather than the chat your colleagues do, with fallbacks, caching and per-request logs. nexos.ai documents budgets and hard caps by user, team or project alongside that12.
Best for: Teams routing production traffic rather than employee chat.
Strengths
Trade-offs
Engineers write the classifier and the rules, so the objective, the eligible models and the logging are all yours. Research routers such as RouteLLM learn the decision from preference data rather than from a hand-written rule34.
Best for: Product teams with measurable high-volume traffic.
Strengths
Trade-offs
A launch brief with the routine stage routed to a cheap model and the judgment stages named by a person, all reading the same project context.
The routing decision is visible at every arrow, because a hidden route makes a quality problem impossible to diagnose. A person checks the extracted fields before the work moves on and reruns the stage on a stronger model when it is thin. The approved brief is then saved with the decisions and the rejected claims, so the next teammate does not repeat the briefing.
Routing adds one requirement to everything else memory has to do, which is that each eligible model receives the same project context before anyone compares them.
A context window is what a model processes in one request, and it fills with prompts, earlier messages and attachments. Memory is stored outside that window and retrieved in later sessions. Chat history is stored text that nothing retrieves for you, so history alone does not carry a project.
Some keep history alone. Some retrieve from documents somebody added. Some learn automatically and keep it to one account, which TeamAI documents and states is not shared. Some store it where the team retrieves it, which WorkLLM documents across project and organisation levels38, 7.
Permissions differ sharply here. TeamAI has private and shared folders where a subfolder inherits its parent's visibility, and Aymo offers viewer, writer and manager access. Several products never say whether a person can inspect every saved entry or keep one session out of it38, 39.
A cheaper model can look worse simply because it received less of the project, so test retrieval separately from model quality. Run one prompt through two eligible models with the same project attached, and check what each was given before judging the router on its answers.
A vendor-neutral plan that starts with a constrained pilot rather than a company-wide Auto default. It takes about two weeks.
Collect a real week of work and split it into routine tasks such as extraction, rewriting, classification and summarising, and hard tasks such as reasoning, analysis and final drafts. The share of routine work is what decides whether routing is worth turning on at all.
Choose whether the router should favour cost, quality or a balance, then give it a short approved model list and exclude anything ruled out by region, security or capability. Keep one named model for regulated or high-stakes work rather than letting the router near it.
Send the same prompts through the router and through your current default model, using real work rather than demo prompts. Check that every eligible model receives the same retrieved project context, because a cheaper model looks worse when it is simply given less.
Track how often somebody reran a task on a stronger model, how many manual edits each route needed, and the time to a useful first answer. Then read the model consumption for both routes, because the saving only counts after the reruns are added back.
Confirm that the selected model is visible on the answer and in the logs, that fallbacks and retries are recorded, and that usage can be read by person, project, model and period. Then check that spending can be capped before a request runs, and read the training and residency terms for every model on the approved list.
Automatic routing is worth turning on when a real share of the work is extraction, rewriting, classification or simple drafting, and when somebody can see which model answered. It saves little when almost every prompt already needs a frontier model. The pattern that holds up is a short approved model list, a stated objective, one named model for regulated work, and a control that reruns a task on something stronger.
Four limits apply. A published saving belongs to somebody else's workload, so Amazon's up to 30 per cent and Microsoft's quality ranges are starting points rather than forecasts. Routing costs predictability, and Microsoft's own limit on the effective context window shows how that becomes a failure rather than a preference. Model prices and model line-ups move, so a router has to be re-measured. And a saving that arrives with more reruns is not a saving.
So the question is what share of your prompts are genuinely routine, and whether anyone would notice a bad route in time to fix it. Measure both on your own work for two weeks with the selected model visible, then decide how much of the queue the router should be allowed to touch.
Upgrade as needed, and only pay for what you actually use
Save ~17% with the annual plan
Pro
Perfect for small and medium teams
Unlimited users & infinite memory
Multi-LLM chats
Granular access control to models
EU data residency
Ultra
Best for large, growing teams
Unlimited users & infinite memory
Multi-LLM chats
Unlimited use of DeepSeek V4 Flash
Granular access control to models
Choose US or EU data residency
Enterprise
Get in touch
For organizations with advanced needs
Unlimited users & SSO
Priority Support
Unlimited use of DeepSeek V4 Flash
Granular access control to models
Choose US or EU data residency
30-days money back guarantee
Playgram will automatically choose the most cost-efficient model suitable for the task. It will be chosen by users in approximately 80% of requests. Your models for the remaining 20%:
If you bought each separately: