AI-native advertising
Advertising built for AI.
Generative AI creates moments that pages, feeds, and search never had: generations, token streams, compute jobs, and visible wait time. wavebird is building advertising formats and measurement models around those moments while keeping ads separate from the model output.
What are inference-time ads?
Inference-time ads are ads shown during the existing wait while an AI model generates its response. They stay separate from the model output and turn a native AI moment into a potential advertising placement.
Fund the compute. Don’t shape the answer.
AI products create value through model output. Advertising should help fund that compute without becoming part of what the model says. wavebird keeps the advertising path separate and builds new formats around the moments AI products already create.
Inference-time ads are a new kind of inventory.
When an AI model generates an answer, the user is already waiting for compute to finish. That creates a natural advertising moment that does not exist in a traditional page view.
An inference-time ad uses that existing generation window. The ad can appear while the model is working, while the AI response remains on its own path.
- 1
User request
The user starts an AI task.
- 2
AI generation begins
The model starts producing the response.
- 3
Ad appears during the existing wait
The separate advertising surface uses the wait already present.
- 4
AI answer completes
The model output completes on its own path.
The wait already exists. The ad does not need to create it.
AI needs formats built for AI.
Inference-time advertising is one format, not the whole category. AI products create several moments where sponsorship can be useful without becoming part of the answer.
- Monetize an existing generation window without inserting advertising into the model output.
- another generation
- additional tokens
- more image generations
- temporary access to a premium model
- longer context
- higher-resolution output
- The ad can be relevant to the moment without becoming part of the model’s words.
- This creates a direct relationship between advertising revenue and the compute that powers the product.
Why ads shouldn’t look like AI answers.
AI products earn trust through their output. Advertising should not borrow that trust by looking like another sentence from the model.
A text ad placed too close to an AI answer can be mistaken for a recommendation, conclusion, or factual statement generated by the product.
Visual advertising has its own language: brand, color, imagery, product, motion, and identity. It can be clearly recognizable as advertising without pretending to be part of the answer.
AI owns the words. The advertiser gets a clearly separate visual surface.
Users should always be able to tell what comes from the AI product and what is sponsored.
Compute Sponsoring connects the ad to the compute moment.
Compute Sponsoring defines a way to associate a sponsorship exposure with a specific unit of computation while keeping the advertising channel separate from model I/O.
The published Compute Sponsoring working paper defines a compute unit as an identifiable unit of active compute. Depending on the system, supported examples are:
- Request
- Generation
- Token block
- Job
That creates something traditional display advertising does not naturally provide: a measurable relationship between an advertising exposure and the compute moment it helped fund.
New inventory needs new metrics.
CPM and CPC remain useful advertising metrics. But AI products also create units that publishers already measure in their products: generations, tokens, compute cost, and time spent waiting for model output.
A measurement vocabulary can connect advertising with the units AI products already measure.
A measurement vocabulary for AI-native advertising.
CPG
Cost Per GenerationThe advertiser cost associated with sponsoring one eligible AI generation.- A generation maps directly to an action the user understands and to a unit of compute the publisher already measures.
CPG = sponsor spend ÷ sponsored generations- Tokens connect advertising economics directly to one of the underlying units of AI compute.
CPT = sponsor spend ÷ sponsored tokensCPT 1K = sponsor spend ÷ sponsored tokens × 1,000CPW = sponsor spend ÷ sponsored wait windowsPublisher metrics should start with the product economics.
Advertisers need buying metrics. Publishers need a different answer: how much of the cost of running the AI product can advertising actually support?
Compute Coverage
Share of compute cost funded by advertisingThe share of AI compute cost funded by advertising revenue.Compute Coverage = publisher ad revenue ÷ AI compute cost- RPG lets publishers compare the revenue generated by AI usage with the cost of serving that usage.
RPG = publisher ad revenue ÷ eligible generations- RPT can be compared with token-level compute cost to understand how advertising affects unit economics.
RPT = publisher ad revenue ÷ generated tokens × 1,000- This shows how much eligible AI usage actually participates in the sponsorship model.
Sponsored Generation Rate = sponsored generations ÷ eligible generationsCompute Coverage example
At 70% Compute Coverage, advertising funds 70% of the measured compute cost in this example.
Above 100%, advertising revenue exceeds the measured compute cost for that scope.
Measure the ad and the compute together.
- generation
- token volume
- compute job
- duration
- compute cost
- eligible opportunity
- filled placement
- impression
- click or interaction
- sponsored exposure
- advertiser spend
- publisher revenue
- CPG / CPT / CPW
- RPG / RPT
- Compute Coverage
Traditional advertising tells a publisher where an ad appeared. AI-native measurement can also show which compute moment the sponsorship was associated with and how that changed the product economics.
From impression reporting to compute economics.
Generation record
Illustrative example only. Values do not represent a guaranteed or current wavebird benchmark.
A new model for advertising in AI.
We believe AI advertising should be designed around the economics and product experience of AI itself, not copied from pages and feeds.
That means formats built around generation moments, clear separation between advertising and model output, and measurement that connects sponsorship with the compute it helps fund.
Inference-time ads are one part of that model. Rewarded compute, sponsored generations, and AI-native economic metrics extend the same idea: advertising can help fund AI without becoming the AI.
Advertising can fund the intelligence without becoming part of the intelligence.
On this page
- ads remain separate from model output
- compute units are explicitly attributable
- new metrics are wavebird proposals
Build advertising around the AI product.
Understand the operating model and publisher economics before choosing an integration path.