Your customer asks your AI product to complete one task. Behind it, OpenAI may classify the request, Anthropic may generate the answer, and another service may retrieve supporting information. If the first attempt fails, part of the sequence may run again.
The customer sees one result. Your business absorbs every call that produced it.
This is why LLM cost tracking cannot stop at checking separate provider dashboards. Those dashboards tell you how much you spent with OpenAI or Anthropic. They do not automatically tell you which customer or workflow created the cost, what produced a spike, or whether the amount you charged covered it.
Multi-provider LLM cost tracking brings usage and cost data from every model provider into one view, then connects that spend to the agent, workflow, customer and outcome that caused it.
Why Multi-Model AI Products Make Cost Tracking Harder
One customer action may trigger classification, retrieval, generation, validation and a fallback call across several models. Input, output, cached and reasoning usage may also be priced differently. OpenAI's pricing documentation, Anthropic's pricing documentation and Google's token-counting guidance show how quickly one "token cost" becomes several cost categories.
The provider may know which account, project or API key made a request, but not:
- Which customer initiated it
- Which agent or workflow they used
- How many attempts were needed
- Whether the task produced a usable result
- What the customer was charged
Your provider bill answers, "What did we spend?" Your internal AI cost tracking must answer, "What did we spend it on?"
What Cost Data AI Companies Need From Each Model Provider
You do not need to collect every field a provider exposes. Start with the information required to explain a cost and connect it to your product.
| Data to Capture | Why It Matters |
|---|---|
| Provider and exact model | Shows where the cost originated |
| Input, output and cached usage | Prevents different usage types from being priced as one |
| Request time and identifier | Connects provider usage with your internal event |
| Estimated request cost | Gives product teams an immediate operating number |
| Workflow and customer ID | Attributes cost to what you sell and who used it |
| Completion and retry status | Includes failed and repeated work |
This is the practical foundation of token usage tracking. Provider prices can change, so use the rate that applied when the usage occurred rather than recalculating historical costs with a new rate.
How to Normalize Token, Request and Model-Level Costs Across Providers
Normalization does not mean pretending every provider measures usage identically. It means organising different provider records so they can be compared at the level your business needs.
| Cost Level | What to Combine |
|---|---|
| Token or usage cost | Input, output, cached and other chargeable usage |
| Request cost | All chargeable usage within one model call |
| Workflow cost | Every request, retry and tool used to complete one task |
| Customer cost | All workflow costs generated by one customer |
| Outcome cost | Total workflow cost divided by completed outcomes |
Keep the original usage record and map it into these shared levels. This helps investigate usage data discrepancies while connecting an OpenAI request, an Anthropic retry and a third-party tool call to the same customer task.
How to Attribute LLM Costs to Agents, Workflows and Customers
Customers do not buy isolated model calls. They buy completed work.
An AI research agent may classify a request, search for information, summarise several sources and generate a final answer. Looking at each call separately hides the cost of the complete workflow.
Every call should therefore carry three identifiers:
- Agent: Which AI product or capability performed the work?
- Workflow: Which complete task did the call support?
- Customer: Which account or tenant requested it?
AWS similarly recommends tracking agentic costs beyond account-level spend. Customer attribution is especially important because two accounts on the same plan can create very different costs. It shows actual cost-to-serve instead of spreading the monthly model bill evenly across every account.
A Worked Example: How One Retry Changes Workflow Margin
Consider an AI support product that charges for each resolved issue. One workflow uses OpenAI for classification, a retrieval service for account information and Anthropic for the response. The first response fails a quality check, so the generation step runs again.
Illustrative scenario: The rates and amounts below are simplified assumptions. They are not current provider prices or Revinci customer data.
| Workflow Component | Illustrative Calculation | Cost |
|---|---|---|
| OpenAI classification | 1 call × $0.004 | $0.004 |
| Retrieval service | 2 searches × $0.010 | $0.020 |
| Anthropic generation | 2 attempts × $0.060 | $0.120 |
| Total workflow cost | $0.144 | |
| Revenue per resolved issue | $0.500 | |
| Contribution after tracked AI cost | $0.500 − $0.144 | $0.356 |
| Illustrative workflow margin | $0.356 ÷ $0.500 | 71.2% |
If the company counted only the successful generation attempt, it would record a cost of $0.084 and estimate an 83.2% margin. Including the failed attempt lowers the margin to 71.2%.
One missed retry overstates the workflow margin by 12 percentage points.
This is the difference between monitoring tokens and understanding the economics of what the customer bought.
How to Set Up LLM Cost Tracking Without Overengineering It
The right setup depends on the stage of your AI product.
| Business Stage | What to Track |
|---|---|
| Testing the product | Provider, model, usage and estimated request cost |
| First paying customers | Add workflow, customer, outcome and retry status |
| Scaling across accounts | Add provider-reported cost, revenue, margin and alerts |
Start with the provider, model, usage, timestamp, request or trace ID, workflow ID, customer ID and completion status. These identifiers let usage roll up into customer-level cost without a complicated finance system.
The amount calculated after a request is an estimate. The provider-reported amount may differ because of credits, discounts, rate changes or missing usage. Keep both figures visible and compare their monthly totals. OpenAI, for example, provides organisation-level reporting through its Costs API.
The boundaries also matter. AI usage tracking records what customers and agents consumed. AI cost observability shows how that consumption affects cost and margin. This article focuses on connecting those two views across several model providers.
How to Compare Model Cost With Customer Revenue
Knowing what a workflow costs is useful. Comparing that cost with what the customer paid makes it commercially useful.
A high-cost workflow may still create enough value to support healthy margin. A low-cost workflow can lose money when bundled into an unlimited plan or used more often than expected.
Compare cost and revenue at the same level. If customers are billed per resolved ticket, measure cost per resolved ticket. If they pay for workflow runs, credits or completed analyses, use that same unit for the cost comparison.
This reveals whether you should:
- Improve an inefficient workflow
- Change an included allowance
- Add an overage
- Price an expensive outcome separately
- Keep the cost because it produces sufficient value
LLM cost tracking and AI billing serve different but connected purposes. Cost tracking explains what the workflow consumed. Billing determines how that usage or outcome becomes a customer charge. Billing accuracy checks should confirm that both sides refer to the same customer, period and commercial unit.
How to Identify Cost Spikes and Margin Leakage Across Providers
A rising provider bill may reflect customer growth. The useful question is whether cost is rising faster than completed work or earned revenue.
Watch for these patterns:
| Signal | What It May Indicate |
|---|---|
| Provider spend rises while completed outcomes stay flat | Retries, loops or model-routing problems |
| Cost per workflow increases suddenly | A model change, longer context or extra tool calls |
| One customer's cost grows faster than its revenue | An allowance or pricing problem |
| Provider cost exceeds internally attributed cost | Missing events or usage data discrepancies |
| Billed usage is lower than completed billable work | Metering gaps and possible revenue leakage |
These signals support revenue leakage prevention by showing where the company may be absorbing cost without capturing corresponding revenue. Before moving to a cheaper model, identify whether the change comes from customer growth, workflow failure, incorrect attribution or a usage-to-billing mismatch.
For an AI company, cost sits inside a larger chain:
Product → Price → Usage → Cost → Margin → Revenue
At Revinci, we connect these stages so teams can see how model consumption affects what they bill and the margin they retain. The value lies in keeping the stages connected as usage and pricing change.
Conclusion: Turn Multi-Model Cost Data Into Better Pricing and Margin Decisions
Using several LLM providers gives your product flexibility. It also spreads the cost of one customer outcome across different models, tools and reports.
Do not begin by measuring every possible field. Start with the five relationships that matter: provider, workflow, customer, outcome and revenue.
When these relationships are visible, you can catch unexplained cost spikes, correct unsustainable allowances and understand customer profitability before month-end.
That is the purpose of multi-provider LLM cost tracking. It is not simply knowing what you spent. It is knowing whether the AI product you are selling can grow profitably.
Frequently Asked Questions About LLM Cost Tracking
How do you track LLM costs across OpenAI and Anthropic?
Collect usage from each provider and attach consistent workflow and customer identifiers to every request. Calculate an estimated cost using the relevant model rates, then compare it with provider-reported cost.
What should an AI company track for every LLM request?
Track the provider, model, input and output usage, timestamp, request or trace ID, workflow, customer, completion status and estimated cost. Add retries, outcomes, revenue and margin as the product reaches paid scale.
What is the difference between token usage tracking and AI cost tracking?
Token usage tracking measures consumption. AI cost tracking applies the relevant model rates and attributes that cost to requests, workflows and customers. Cost tracking is therefore the more commercially useful view.
Should failed LLM calls and retries be included in cost?
Yes. Failed calls, retries and fallback models still create provider costs. Include them in the cost of the completed workflow even when the customer is billed only once.
Can LLM cost tracking show AI product profitability?
Not by itself. Profitability requires the complete workflow cost to be connected to the customer or billable outcome and compared with the revenue it generates.