Database Health
Deep operational monitoring for managed databases — connections, storage, latency, replication lag, query performance, forecasting, and executable auto-actions.
At a glance
| Route | /database-health |
| Group | Analyze |
| Page permission | intelligence |
What it is
Deep operational monitoring for managed databases — principally AWS RDS and Aurora, with Azure and GCP pull paths — covering connections, storage, latency, replication lag, query performance, rule-driven recommendations, alerting, and Autopilot-executable auto-actions.
Who it is for
Database administrators, platform engineers, and SREs responsible for database reliability and cost — and FinOps practitioners who have found that databases are frequently the largest and least examined line on the bill.
How it works
Metric history you would not otherwise have
Metrics are pulled per instance and persisted, giving a history that provider consoles typically do not retain at useful granularity.
The direct connection unlocks the best findings
An optional direct database connection — credentials encrypted at rest — enables introspection and query-level sampling that cloud metrics alone cannot provide.
This is where the highest-value findings come from. Cloud metrics tell you an instance is busy; query-level sampling tells you which query is making it busy.
Features
- Per-instance monitoring of connections, storage, latency, and replication lag
- Metric history retained beyond provider console retention
- Query-level sampling and a top-queries view where a direct connection is configured
- Schema and configuration introspection
- Rule-driven recommendations with per-instance findings
- Storage and capacity forecasting, so exhaustion is anticipated rather than discovered
- Log-based diagnostics reading database logs for problems metrics do not expose
- Configurable alert rules with notification
- Autopilot-executable auto-actions on database findings
- Azure and GCP pull paths alongside AWS
- A demo data mode for evaluation without a connected database
- Manual sync trigger alongside scheduled collection
How to use it
Connect the account hosting the databases and sync
To populate the instance list.
Confirm every production database appears
In the inventory. A missing instance is a blind spot you will not otherwise notice.
Configure a direct connection for your most critical instances
To unlock query-level sampling and introspection. This is where the highest-value findings come from.
Work the findings per instance
Cost items: over-provisioned instances and unused replicas are typically the largest.
Reliability items: replication lag and connection saturation are the largest.
Act on storage forecasts before exhaustion
Not after. That is the entire point of forecasting capacity.
Configure alert rules on the metrics that matter
And route them to your workspace's alert channel in Integrations.
Why it matters
Databases are simultaneously among the most expensive and most operationally sensitive resources in a cloud estate — which is exactly why they are usually left alone. Nobody wants to rightsize a production database on a hunch.
By combining measured metric history, query-level evidence, and capacity forecasting, this module makes database optimisation a defensible decision rather than a risk.
And the reliability findings prevent outages that cost far more than the infrastructure.
Connects to
- Findings flow to Opportunity Queue
- Can be executed via Autopilot and Optimization
- Credentials managed alongside Cloud Accounts
- Alerts route through Integrations