AI Infrastructure
A dedicated cost and waste surface for machine-learning and generative-AI infrastructure, with real findings and one-click actions rather than generic compute analysis.
At a glance
| Route | /ai-infrastructure |
| Group | Operate |
| Page permission | overview |
What it is
A dedicated cost and waste surface for machine-learning and generative-AI infrastructure: Amazon SageMaker and Bedrock, Google Vertex AI and Gemini Enterprise, and Azure OpenAI and AI Foundry.
It produces real findings and one-click actions rather than generic compute analysis.
Who it is for
ML platform teams, AI engineering leads, and the FinOps practitioners who have discovered that AI infrastructure has become one of the fastest-growing and least-governed lines on the cloud bill.
How it works
Why AI workloads need their own detectors
General compute detectors do not serve AI workloads well:
- GPU instances idle differently than conventional compute
- Endpoints are provisioned differently — a hosted endpoint bills whether or not it serves traffic
- Inference cost scales with token throughput, not instance hours
The platform therefore maintains provider-specific AI signal modules that inspect these services on their own terms.
An AI service taxonomy classifies spend across training, inference, endpoint hosting, and model API consumption — so the bill can be reasoned about by workload type rather than by SKU.
Findings carry the same savings, confidence, effort, and risk scoring as every other detector output, and each supported provider has its own AI remediation action module for one-click execution.
Features
- Findings and one-click actions across SageMaker, Bedrock, Vertex AI, Gemini Enterprise, Azure OpenAI, and Azure AI Foundry
- AI spend trend and anomaly detection specific to AI workload shapes
- AI-specific recommendations with accept and dismiss workflow
- Configurable AI cost guardrails, per workspace
- Model pricing catalogue and a side-by-side model cost comparison tool
- Contextual analysis linking AI spend to the workloads generating it
- Direct ingestion endpoint for AI cost data from sources outside the standard billing path
- GPU intelligence — per-container GPU cost attribution and utilisation efficiency (Business and above)
How to use it
Connect the accounts hosting AI workloads
And confirm ingestion covers the relevant services.
Identify what dominates
Review the AI spend trend and determine which of training, inference, or endpoint hosting is the largest share. The right optimisation differs completely between them.
Work the findings list
Idle endpoints and over-provisioned GPU instances are typically the largest and safest early wins.
Test cheaper models on high-volume, low-complexity calls
Use the model comparison tool to check whether a cheaper model meets your quality bar. This is often the single largest lever available.
Set guardrails on the workloads that matter
So AI spend growth raises an alert rather than a surprise invoice.
Accept and dismiss deliberately
Accept the recommendations to act on; dismiss those that are intentional. A queue full of known-intentional items stops being read.
Connect direct LLM vendor spend
In AI Provider Accounts, so consolidated AI spend is actually complete.
Why it matters
AI infrastructure is where cloud spend is growing fastest and where governance is weakest — largely because general-purpose cost tools cannot see it properly.
An idle GPU endpoint can cost more per month than an entire conventional application estate.
This module gives that spend the same detection, action, and guardrail treatment the rest of the estate already receives.
Connects to
- Complemented by AI Provider Accounts for direct LLM vendor spend
- Distinct from Membership → Apex AI Credits, which meters the platform's own AI consumption
- Findings flow to Opportunity Queue and Optimization
- Unit cost per inference is tracked in Unit Economics
Optimization
The action-first savings pipeline — bundles opportunities into executable work, runs them through guardrails and approvals, and reports realised return on investment.
CEO View
A deliberately minimal executive snapshot — the biggest waste source, the ROI narrative, and the one action that should happen next.