AI is cheap to start and surprisingly easy to overspend on. Costs are usage-based rather than a fixed subscription, they're spiky by nature since usage varies week to week, and they're largely invisible until the invoice arrives, because most tools give you a single monthly total rather than a breakdown of where it came from. Here is a practical way to keep token spend predictable as adoption grows, without slowing your teams down or rationing access to the tools that make them faster.
Understand what a token actually costs
Providers bill per token of input and output, and prices vary widely between models, sometimes by an order of magnitude for tasks that look similar on the surface. A frontier-tier model can cost many times what a smaller model does for the exact same request, and long prompts, large pasted documents, and verbose outputs all add tokens on top of that base rate. The first step toward controlling spend isn't a policy or a budget; it's simply being able to see cost per request, broken out by model, so you know where the money is actually going before you try to change anything.
Attribute every request to a team
A single monthly number tells you almost nothing you can act on. When every request is tied to the person and team that made it, patterns emerge quickly: which teams are heavy users, which models they gravitate toward, and where cost is clearly outrunning the value of the work being done. Attribution is also what makes accountability possible without micromanaging; a team that can see its own spend trend tends to self-correct faster than one waiting for a top-down mandate.
Route cheap work to cheap models
Most requests, honestly, don't need your most expensive model. Formatting a document, summarizing a short email, doing a quick lookup, or rewriting a paragraph are all tasks a smaller, faster model handles just as well as a frontier one, at a fraction of the cost. Reserving frontier-tier models for genuinely hard reasoning, long-form drafting, or high-stakes analysis, and routing everything else to a cheaper model by default, is the single biggest lever on overall cost. Done well, it can cut spend substantially with no visible drop in the quality of everyday work, because the routine tasks were never where the expensive model's extra capability was being used anyway.
Set budgets before the invoice, not after
Reactive cost control, noticing a spike after the bill lands, always costs more in cleanup than a small amount of upfront structure would have. A few habits make the difference:
- Give each team its own budget and an alert as it approaches the limit, rather than one company-wide number nobody owns.
- Watch cost-per-team trends weekly, not monthly; a spike caught in week one is a quick conversation, while the same spike caught at month-end is already baked into the invoice.
- Default every user to a sensible, cost-appropriate model and let them escalate to a stronger one deliberately when a task calls for it, rather than defaulting everyone to the most expensive option.
- Kill runaway usage patterns, like a script or workflow looping on an expensive model, while they're still small rather than after they've run for a full billing cycle.
Watch for the hidden multipliers
Two patterns quietly multiply spend beyond what any single request would suggest. First, long conversation threads resend their entire history on every turn, so a chat that's been running for an hour can cost far more per message than it did at the start, even though each individual question looks small. Second, pasting large documents repeatedly across separate questions means paying to resend that document every time, rather than once. Neither shows up as an obvious red flag in a monthly total; both show up clearly once requests are broken out individually.
How Switchboard helps
Switchboard does all of this out of the box: every request is metered and attributed to a team automatically, budgets are enforced per team rather than tracked after the fact, and routing sends work to an appropriate model by default so nobody has to manually pick the cheapest option for routine tasks. Token costs are passed through at provider rates with no markup, so what you see is what the provider actually charged. You get one live view of AI spend across every provider and every team, plus the guardrails to keep that spend predictable as adoption grows rather than discovering the problem in an invoice.