The economics of enterprise AI just changed in public. Financial Times reporting published September 27 says open-weight models now account for 56% of Vercel’s AI Gateway token volume, up from 7% in December, while open models handle 40% of AT&T’s AI workloads. Executives mentioned “open-weight” or “open-source” models six times more often on U.S. earnings calls than a year ago. AT&T says routing selected work to less expensive models cut coding costs by as much as 56%; Coinbase reported savings of about 50%. (Financial Times; AI Weekly summary)

For marketing leaders, this is not a story about replacing every premium model. It is a signal that the default architecture for AI-powered marketing is becoming routing: use the least expensive model that clears the quality bar, and reserve frontier systems for work where a mistake is costly.

Token volume is moving faster than AI budgets

Vercel’s official September AI Gateway Production Index, based on data through August 2026, shows open-weight models crossing a meaningful threshold: they processed 56% of gateway tokens but represented only 14% of gateway spend. Their share rose every month from 13% in April to a majority in August. Vercel also reported that average price per token fell 23.2% in August, the third consecutive monthly decline. (Vercel AI Gateway Production Index)

That gap between volume and dollars is the strategic point. Routine marketing work generates enormous token volume: classifying leads, summarizing calls, tagging creative, clustering search queries, drafting variants, transforming product feeds and producing internal reporting. A company does not need its most capable reasoning model for every one of those tasks. It needs reliable output, predictable latency and a cost profile that allows the team to run the workflow continuously.

Model routing turns AI from a tool bill into an operating system

AT&T offers the clearest enterprise pattern. Its Ask AT&T platform processes roughly 45 billion tokens per day for about 100,000 employees. Using LiteLLM to route low-complexity work to models such as NVIDIA Nemotron, Meta Llama and Google Gemma, the company sends harder work to premium GPT and Claude systems while keeping routine requests on lower-cost models. AT&T’s Mark Austin says the company is targeting 60% to 70% open-model usage within a few years. (AI Weekly report on AT&T’s routing strategy)

The mechanism is increasingly accessible. LiteLLM documents routing strategies based on capacity, latency, fallbacks and lowest cost, while Vercel AI Gateway offers provider ordering, model fallbacks, request logs, token tracking and spend budgets. (LiteLLM routing documentation; Vercel AI Gateway documentation) The marketing implication is practical: teams can design a service tier for each workflow instead of forcing every prompt through one expensive vendor.

Quality gates matter more than model labels

Open-weight adoption does not mean “open is always better.” Stanford’s 2026 AI Index found that, as of March, the top closed model led the top open model by 3.3% on its cited Arena comparison, and six of the top ten models were closed. The same report says competitive pressure is shifting toward cost, reliability and domain-specific performance. (Stanford HAI 2026 AI Index)

That is exactly why marketing organizations need evaluation gates. A low-cost model can handle a product-description rewrite or intent label if it passes factuality and brand checks. A premium model, human review or deterministic rule should handle regulated claims, executive messaging, crisis communications, customer-specific recommendations and anything that can create legal exposure. The right question is not “Which model won?” It is “Which model passes this workflow’s acceptance test at the lowest total cost?”

What marketing leaders should do in the next 30 days

  1. Inventory by task, not department. Measure tokens, latency, error rates and human-review time for content, paid media, CRM, analytics and customer-service workflows.
  2. Create three quality tiers. Route repetitive transformations to an economical model, analysis and generation to a mid-tier model, and high-stakes work to a frontier model or human approval.
  3. Run a controlled routing pilot. Select one high-volume workflow, hold prompts and inputs constant, and compare cost per accepted output—not cost per token alone.
  4. Keep a model ledger. Record provider, model version, license, data-retention terms, evaluation scores, spend and fallback behavior so a model change does not silently change brand risk.

The competitive advantage will not come from proclaiming loyalty to one AI lab. It will come from building a measurable system that can switch models without sacrificing accuracy, brand consistency or accountability. The companies learning that discipline now will be able to run more experiments, personalize more customer journeys and improve AI-search visibility without allowing inference costs to consume the margin.

Need a practical AI marketing architecture that balances performance, visibility and cost? Real Internet Sales helps businesses build and measure AI-powered marketing systems. Call 803-708-5514 or visit realinternetsales.com.