Large language models (LLMs) are the “brains” of AI and they’re driving a fundamental shift in the economics of IT. Using new AI-specific model architectures, plus big data and multibillion parameter modeling, LLMs are set to be the foundation of a new AI tech stack. Advances in GPU hardware and chip innovation are also critical components – but building or consuming a modern AI stack comes at a cost.
Shifting the cost paradigm: from time to tokens
Traditionally, AI costs were driven by CPU-centric virtual machines (VMs) and based on usage time. But that cost metric doesn’t capture how many of today’s AI models are billing back to enterprises. Enter the token.
A token represents a small unit of data – for example, part of a word, image, audio clip, or video. During inference or model training, AI models break down inputs into these units – or tokens – and process them step by step. The number of tokens used in an activity is now the primary billing metric in most AI services. Organizations pay by the number of tokens their AI consumes – whether that’s in prompts, outputs, or intermediate reasoning steps.
One caveat: Token-based AI pricing introduces unpredictability because costs vary with prompt length, complexity, usage, and more. Unlike with traditional cost models, this can make budgeting challenging, and it can complicate planning for AI initiatives.
The paradox of lower token costs and higher usage
Token costs have dropped, with substantial reductions over the last two years. Activities that once cost dollars-per-thousand tokens now cost fractions of a dollar-per-million. That should mean savings, right?
Well, not exactly. This is where Jevons’ Paradox comes in. As the cost per token falls, organizations don’t necessarily spend less; they just consume more. Complex tasks that were once too expensive to justify are now within reach. Chain-of-thought reasoning, for example, enables models to generate more accurate and nuanced responses, but it can consume up to 100-times more tokens per inference than traditional reasoning tasks. This trend will continue, especially as AI becomes embedded in more workflows.
Tokens are the new reality of AI strategy
Today tokens are the basis for most commercial model provider pricing plans. Although token costs may not have been a critical factor during proof-of-concept initiatives, as organizations expand their AI use, leaders should be keenly aware that token costs will likely shape decisions about everything from the tech stack to vendor partnerships.
The decision that many are facing as they scale on whether to continue to pay for off-premises proprietary models or to build an AI factory onto self-managed infrastructure is becoming increasingly important. Open-source models, combined with best-in-class AI factory stacks, might avoid per-token fees but require more investment in the tech stack, whereas proprietary models via APIs often charge per token but can be much faster to deploy and easier to manage. The economics vary sharply depending on scale, sensitivity, and predictability of demand.
The answer is often to embrace a fit-for-purpose mentality leading to a hybrid AI approach. AI factories become logical when scale and predictability cross a certain threshold leading to better economics.
The tradeoffs are real, and they can be difficult to make. However, leaders should assume that tokens represent the future basis of fundamental AI cost calculations and that token usage will continue to grow – perhaps exponentially. Their IT and business strategy should reflect this new reality.
Aligning token use and business strategy
As the foundation for measuring AI costs, token use must be managed with the same rigor as cloud spend or hardware capacity. And like any other resource, token consumption should be aligned with business priorities.
Some workloads require high token use to deliver value. For example, a customer support agent that needs detailed reasoning may justify the expense. But other use cases, like summarization or document classification, may be better served by smaller models that can be deployed with fewer tokens.
The key is to map token use to business impact. This means selecting the right model for each task, choosing the right hosting strategy (on-premises, hyperscaler, or hybrid), and optimizing user experience to reduce waste like bloated responses or unnecessary processing.
Monitoring token usage in real time is also essential. Time to first token, latency between tokens, and total tokens used per interaction are all critical metrics to optimize performance and control costs.
The takeaway
A discussion of tokens may seem overly technical, but it’s not when they lie at the heart of the economics of AI. They show you exactly what you're paying for – and how much – and whether that spend is delivering the value you seek.
For tech and business leaders alike, it’s clear: Tokens are not just a technical metric, they're a key strategic lever for managing AI at scale. Monitor their consumption, plan around them, and optimize their use. Because in the new AI economy, tokens are the coin of the realm.
Comments