Knowing how much you spent on AI is useful.
Knowing which customer created that cost and whether the revenue from that customer still leaves enough margin is much more useful.
An AI company may know it spent $100,000 across models, tools and infrastructure last month and still struggle to answer:
- Which customers cost the most to serve?
- Which agents or workflows are becoming more expensive?
- How much do retries and failed attempts add?
- Does the pricing model still work after delivery cost?
- Which customers are actually profitable?
That is where AI cost observability needs to go beyond an LLM-spend dashboard.
Knowing what an AI workload costs is only half the answer. The commercial question is what that workload cost for a specific customer relative to what the customer paid.
What Is AI Cost Observability?
AI cost observability is the ability to trace AI delivery costs from models, tools and infrastructure to the agents, workflows and customers that created them.
Basic cost monitoring might tell you:
We spent $25,000 on Model A.
Cost observability should help you move toward:
Customer A created $3,400 of AI delivery cost across 8,200 workflows.
And eventually:
Customer A paid $8,000, cost $3,400 to serve, and generated $4,600 in gross margin.
The progression is:
AI Spend → Cost of Work → Customer Cost-to-Serve → Margin
That is when AI cost data becomes commercially useful.
Why LLM Spend Is Not Customer Cost-to-Serve
An LLM bill captures only part of the cost of serving an AI customer.
A single agent workflow can involve:
Model Calls + Tokens + Tools/APIs + Compute + Infrastructure + Retries + Failed Attempts
LLM Spend ≠ Customer Cost-to-Serve
A more useful starting point is:
Customer Cost-to-Serve = Attributable Model Cost + Tool/API Cost + Infrastructure Cost + Other Attributable Delivery Cost
The exact cost categories will vary by product.
The cost needs to be connected to the customer, agent or workflow that created it.
From Spend to Margin: 4 Levels of AI Cost Visibility
1. AI Spend
Question: How much are we spending?
This includes visibility into providers, models and major infrastructure costs.
| Cost Area | Monthly Spend |
|---|---|
| Model Provider A | $42,000 |
| Model Provider B | $21,000 |
| Search / Tool APIs | $9,000 |
| Other AI Infrastructure | $8,000 |
This helps Finance and Engineering understand total spend. But it does not answer:
Who created the cost?
2. Cost of Work
Question: What does the AI actually cost to operate and complete work?
Now the cost is attached to an agent, a workflow, a resolution, an output, or another unit of work.
| Unit of Work | Average Cost |
|---|---|
| Lead researched | $0.42 |
| Support resolution | $0.68 |
| Invoice processed | $0.31 |
| Research report | $1.84 |
This is much closer to a pricing decision.
If a workflow is sold for $1.00 but regularly costs $0.80 to execute, that is a very different commercial problem from a workflow costing $0.20.
3. Customer Cost-to-Serve
Question: Which customer created the cost?
This is where averages begin to become commercially actionable.
Two customers can buy the same product and pay the same amount while creating very different economics.
| Customer A | Customer B | |
|---|---|---|
| Monthly Revenue | $10,000 | $10,000 |
| Workflows | 12,000 | 7,500 |
| Model + Tool Cost | $5,100 | $2,300 |
| Other Attributable Cost | $900 | $500 |
| Customer Cost-to-Serve | $6,000 | $2,800 |
Both customers generate $10,000 in revenue. But Customer A costs more than twice as much to serve.
That difference can affect:
- pricing,
- renewals,
- usage allowances,
- enterprise discounts,
- model routing,
- minimum commitments,
- contract terms.
4. Cost-to-Margin
The final question is:
What did the customer pay, what did serving them cost, and what did we keep?
That creates the full commercial view:
Revenue → Usage → Cost → Margin
At this point, AI cost observability is no longer only an engineering or FinOps exercise.
It becomes part of pricing and revenue management.
Why AI Agent Cost Can Vary So Much
Agent cost is not always predictable at the task level.
An agent can:
- reason for longer,
- select a different model,
- invoke more tools,
- take another workflow path,
- retry,
- fail and try again.
McKinsey's 2026 research on agentic economics cites programming research where separate completions of the same task varied in cost by as much as 30×.
That matters commercially. If you charge $1 per request or $5 per workflow, the revenue may stay fixed while the delivery cost behind that unit changes significantly.
Average cost can hide the customers, workflows and executions where margin is actually breaking.
What Costs Should Be Attributed?
Tokens are important, but they are only one part of the picture.
| Cost Signal | Why It Matters |
|---|---|
| Models & Tokens | Base inference cost |
| Tools / APIs | External execution cost |
| Compute / Infrastructure | Cost of running the workload |
| Retries / Failures | Cost that may create little or no billable value |
| Customer / Agent / Workflow | Shows who or what actually created the cost |
The important word is attribution.
A token event on its own tells you consumption. A token event connected to:
Customer → Agent → Workflow
starts to explain commercial economics.
How to Calculate Customer Cost-to-Serve
Start with all attributable delivery cost associated with a customer. A practical structure is:
Customer Cost-to-Serve = Model Cost + Tool / API Cost + Compute & Infrastructure Cost + Other Attributable Delivery Cost
From there, the cost can be connected to the commercial unit.
Cost per Workflow
Customer Cost-to-Serve ÷ Completed Workflows
Cost per Resolution
Cost of Eligible Attempts ÷ Successful Resolutions
Cost per Agent
Attributable Agent Cost ÷ Active Agents
Customer Cost-to-Serve
Total Attributable AI Delivery Cost for the Customer
The exact allocation method depends on the architecture. The important rule is simpler:
Cost becomes useful when it can be attached to the unit you price, sell or manage.
Illustrative Customer Cost-to-Serve Example
Consider an AI research product. A customer completes 10,000 research workflows in one month.
| Revenue / Cost Layer | Amount |
|---|---|
| Customer Revenue | $10,000 |
| Model Cost | −$2,100 |
| Search & Tool APIs | −$850 |
| Compute & Infrastructure | −$450 |
| Retry / Failure Cost | −$600 |
| Customer Cost-to-Serve | −$4,000 |
| Gross Margin | $6,000 |
| Gross Margin % | 60% |
The numbers are illustrative, not an industry benchmark. But they show the difference between:
We spent $2,100 on models.
and:
This customer cost $4,000 to serve and generated a 60% gross margin.
The second answer is much more useful for pricing.
The Revinci Cost-to-Serve View
At Revinci we look at AI cost through three connected questions:
Who consumed it?
Customer.
What created it?
Agent / Workflow.
What commercial result sits above it?
Revenue / Margin.
That creates a different progression from a traditional model-cost dashboard:
Customer → Agent / Workflow → Cost-to-Serve → Revenue → Margin
The point is not simply to observe AI infrastructure. It is to connect operational cost to the commercial model built on top of it.
How Cost Observability Changes Pricing Decisions
Imagine the research workflow above rises from $0.40 cost per workflow to $0.72 cost per workflow while the customer still pays $1.00.
The system may still be working perfectly. Usage is captured. The invoice is correct. Revenue is being collected. But the economics have changed.
The business may now need to:
- change model routing,
- reduce expensive tool usage,
- optimize the workflow,
- change included usage,
- introduce tiers,
- increase the unit price,
- renegotiate the enterprise contract,
- move to another pricing model.
This is where AI cost observability connects directly to AI agent pricing.
You cannot properly evaluate a pricing model if you cannot see the cost underneath the unit being priced.
How SmartCost + SmartMargin Operationalize It
At Revinci we approach the problem through SmartCost + SmartMargin.
SmartCost — Attribute Cost
Revinci SmartCost is designed to move from aggregate AI spend toward attributable cost-to-serve.
At Revinci we describe cost visibility across:
Customer → Agent → Workflow
and across cost categories including:
Tokens → Compute → Storage → APIs
That helps move the question from:
What did our model provider cost?
to:
What did this customer's agent and workflows cost to serve?
Bill — Connect the Usage Signals
Attribution also depends on understanding what happened.
Revinci Bill includes metering across AI revenue signals including:
- tokens,
- actions,
- tool calls,
- workflows,
- outcomes,
- other usage events.
Those signals provide the operational context needed to connect cost to the commercial unit.
SmartMargin — Connect Cost Back to Revenue
Once the delivery cost is attributable, the next question is margin.
Revinci SmartMargin is designed to connect cost back to revenue across customers, agents and deals.
That closes the loop:
Usage → Cost → Revenue → Margin
The goal is not simply to reduce AI spend. It is to understand whether the revenue model remains economically sustainable.
AI Cost Observability vs AI Margin Intelligence
The two ideas are related but different.
| Capability | Main Question |
|---|---|
| AI Spend Monitoring | How much are we spending? |
| AI Cost Observability | Where did the cost come from? |
| Customer Cost-to-Serve | Which customer created the cost? |
| AI Margin Intelligence | What did that customer pay and what did we keep? |
At Revinci our opportunity sits in connecting all four.
Frequently Asked Questions
What is AI cost observability?
AI cost observability is the ability to trace AI delivery costs from models, tools and infrastructure to the agents, workflows and customers that created them.
What is customer cost-to-serve for AI?
Customer cost-to-serve is the attributable cost required to deliver an AI product to a specific customer. It can include model usage, tools, APIs, compute, infrastructure and other directly attributable delivery costs.
Why is LLM cost tracking alone not enough?
LLM cost tracking explains model spend. It does not automatically show which customer, agent or workflow created the cost, or how that cost compares with customer revenue.
Why can the cost of the same AI task change?
Agents can follow different paths, use different models or tools, generate different amounts of tokens and retry actions. McKinsey cites research showing as much as a 30× cost difference between separate completions of the same agentic programming task.
What is the difference between cost-to-serve and margin?
Cost-to-serve measures the attributable cost of serving a customer. Margin compares that cost with the revenue generated from the customer.
Move From AI Spend to Customer Economics
Knowing your AI spend is the starting point. The commercially useful view is:
AI Spend → Cost of Work → Customer Cost-to-Serve → Revenue → Margin
That is where cost data begins to answer the questions Finance, RevOps and pricing teams actually need answered.
Because the important question is no longer only:
How much is our AI costing us?
It is:
What does each customer cost to serve, and does the revenue model still work?