SemiAnalysis Tokenomics Model

Tokenomics Model

The hardware inputs of AI connected to the software outputs built on top of it. The first demand-side model of AI translates revenue and usage growth into future hardware demand across Nvidia, AMD, Google TPU, and Amazon Trainium fleets, for everyone following the money trail.

Scope
Demand side
Build
Bottoms-up
Workloads
4 profiled
Revenue lines
3 modeled
Illustration of a factory producing crates of charts on a conveyor, trucks hauling them to market, and a hand holding a clipboard of graphs

The multi-trillion dollar question, made calculable.

The SemiAnalysis Tokenomics Model connects the hardware inputs of AI to the software outputs of the services built on top of AI models. It completes our end-to-end coverage of the AI and chip industry with the tools and metrics to calculate the ROI on AI spend, adoption of AI by use case, and the growth of new business models enabled by AI.

Investors, corporates, and policy makers following the money trail of AI can disentangle the economic relationships between the players in the AI value chain and answer the question the whole buildout hangs on: what is the ROI and profitability of AI business models?

Who uses it, and what it decides.

The buyers of this model, and the calls they make with it.

Public & private market investors
Positioning across the AI value chain with a demand-side view: which business models actually earn a return on the compute they buy, and whose revenue holds up as token economics shift.
Corporate strategy & finance teams
ROI on AI spend before the budget commits: adoption by use case, the unit economics of token consumption, and where AI spending actually pays back.
Policy makers & economists
The token economy sized and followed: where AI value accrues, what the money trail funds, and how usage growth turns into infrastructure demand.
Software vendors & AI-native builders
The unit economics of the new software model: what serving costs per token, what the market bears, and how token consumption pricing disrupts the seat-based incumbents.

From installed base to income statement.

The fleet is counted, its token output is modeled from first principles, the token economy is sized, and the loop closes back into hardware demand.

The fleet, counted

A detailed forecast of the AI hardware installed base and future investments, buyer by buyer:

Hyperscalers

Microsoft, Google, Amazon, Meta, and Oracle, with installed base and future investment tracked per buyer

Foundation labs

OpenAI, Anthropic, and DeepSeek as compute demand sources, tracked alongside Thinking Machines and peers

Neoclouds

Coreweave, Nebius, Crusoe, and the quality neocloud supply base as measured by our ClusterMAX ratings

Tokens per second, from first principles

Bottoms-up token throughput forecasts built from the three things that set them:

Hardware systems

GB200 NVL72, VR144, CPX, TPU v7, Trainium 3, and the systems that follow them

Model architectures

GPT 5, Sonnet 4, DeepSeek V3, Kimi K2, and the architecture choices that move throughput

User workloads

Coding, chat, document analysis, and agentic behavior, each with its own token profile

The token economy, sized

Addressable market analysis across everywhere tokens are consumed:

In-app usage

Existing application token usage across Google AI Overviews, ChatGPT, Grok, and Meta AI

API endpoints

Inference endpoints for ChatGPT, Claude, Qwen, DeepSeek, Llama, and the rest of the serving market

Token-native software

Cursor, Windsurf, Harvey, Perplexity, and the emerging startups whose COGS is tokens

The money trail, closed

Usage and revenue growth translate back into aggregate hardware demand: inference demand from adoption, training demand from architecture development and expected returns, ROIC by deployer, and revenue and profit forecasts across rental, model, and software lines, including the disruption of seat-based SaaS by token consumption pricing at companies like Salesforce, Workday, Adobe, SAP, ServiceNow, and Atlassian.

Follow the money trail.

AI economics is a loop. Capex buys compute, compute produces tokens, tokens earn revenue, and the returns decide the next order of hardware. This model is the first to close that loop from the demand side.

The loop · Capex

Fleets get bought

Hyperscalers, foundation labs, and neoclouds put capital into accelerator fleets.

In the model: installed base and investment forecasts, buyer by buyer

The loop · Compute

Fleets get scheduled

GB200 NVL72 to TPU v7 turn capex into serving capacity, set by hardware and architecture.

In the model: bottoms-up tokens per second by system and architecture

The loop · Tokens

Tokens get consumed

Chat, coding, document analysis, and agentic workloads burn tokens through apps, APIs, and token-native software.

In the model: the addressable market of the token economy, channel by channel

The loop · Revenue

Returns close the loop

Rental, model, and software lines split the take, and ROIC decides the next order of hardware.

In the model: revenue, profit, and ROIC by business model, translated into future GPU demand

The ring is a map, not a chart. In the model, every turn of the loop carries numbers: installed base by buyer, tokens per second by system, consumption by channel, and ROIC by business model.

How the model is built.

Supply of compute, production of tokens, and the demand that pays for both, forced to agree.

  1. Count the compute

    Installed base and investment forecasts across hyperscalers, foundation labs, and neoclouds set the supply of AI compute, with quality graded by ClusterMAX.

  2. Model the tokens

    Bottoms-up throughput per hardware system, model architecture, and user workload turns compute into token supply, and usage tracking turns applications into token demand.

  3. Close the loop

    Adoption, revenue, and ROIC translate back into future hardware demand, split into inference and training, so the money trail runs end to end.

Research that ships with the model.

Model subscribers receive the update notes, webinars, and analysis published against each release. A sample of recent coverage:

The archive comes with the model.

Every release ships with notes and webinars like these, written by the analysts who maintain the numbers.

Get access

Common questions.

Anything not covered here, ask the team directly through the form below.

What does the Tokenomics Model include?

Installed base and investment forecasts for hyperscalers, foundation labs, and neoclouds, bottoms-up token throughput by hardware system, model architecture, and workload, addressable market analysis of the token economy, tracking of SaaS disruption by token consumption pricing, ROIC of AI deployments, and aggregate inference and training hardware demand with revenue and profit forecasts across rental, model, and software lines.

How is the model built?

Demand side first. Token consumption is tracked across applications, API endpoints, and token-native software, and set against bottoms-up token production per hardware system and model architecture. Usage and revenue growth then translate into future hardware demand and ROIC by deployer, closing the loop between AI spending and AI income.

How is the model delivered?

As an Excel workbook with dashboard access, including one year of quarterly updates, an onboarding call with the team to walk through the model and methodologies, and ad-hoc calls for questions that come up in use.

Is it part of the SemiAnalysis newsletter subscription?

No. Industry models are separate institutional offerings and are not included with the annual newsletter membership.

Can it translate AI usage into GPU demand?

Yes. It is the first demand-side model of AI, built to translate revenue and usage growth into future hardware demand across Nvidia GPUs, AMD GPUs, Google TPUs, Amazon Trainium, and more, split into inference demand from adoption and training demand from model development.

Models that pair with this one.

Tokenomics is the demand side of a chain the other models cover link by link: the silicon, the facilities, and the cost of running both.

AI Cloud TCO Model

Supply-side twin

The cost of producing a token: what a GPU-hour costs to rent and run, cluster by cluster.

Rental economics Institutional
View model

Accelerator & HBM Model

The silicon

The accelerator shipments that the hardware demand this model forecasts ultimately orders.

SKU coverage Institutional
View model

Datacenter Industry Model

The facilities

Where the fleets serving these tokens get built, powered, and energized, site by site.

Site-level Institutional
View model

Get the Tokenomics Model.

Start with the sales team. They come back with scoping, licensing, and pricing for your mandate.

  • Scoped to your use case
  • Onboarding and ad-hoc analyst calls included
  • Custom research engagements available

Tell us what you are trying to decide. The sales team will follow up to scope coverage and provide pricing.

← Back

Thank you for your response. ✨