Re-evaluated, not re-audited
Scenarios run continuously against current usage rather than during a quarterly review, so a regression is caught in the week it appears.
Domain edition
DigiUsher ingests data platform cost from system tables rather than invoice summaries: DBUs by job and cluster, credits by query and warehouse, slot-seconds by job, Atlas tiers by project.
That grain is what turns a platform total into an attributable cost, and what lets a schema change upstream be connected to the cost regression it caused three jobs downstream.
Available today system-table grain
Coming soon on the connector roadmap
Every platform lands in the same FOCUS schema, so a Databricks DBU, a Snowflake credit and a BigQuery slot-second are directly comparable, and join to the cloud, Kubernetes and AI cost that surrounds them. A new platform is a connector, not a platform release.
A monthly platform total cannot explain why Tuesday’s job costs three times what it cost last month. DigiUsher ingests at the grain where the answer lives, the job, the query, the warehouse, the cluster, and then joins that grain to the pipeline, dataset, team and downstream consumer it belongs to.
The cause of a cost spike is rarely where it surfaces. Attribution without lineage just relocates the argument.
DBUs by job and cluster, credits by query, user and role, slot-seconds by job. The single unpartitioned scan repeated hourly becomes attributable to someone who can fix it.
GrainWarehouse and cluster configuration against observed utilization: auto-suspend thresholds, sizing, all-purpose clusters used for scheduled work, reservation and slot commitments.
GrainTime-travel retention, fail-safe storage, unused clones, staging tables that outlived their migration, and Atlas tiers that were raised once and never lowered.
GrainCost rolled up per pipeline run and per dataset, so a schema change upstream is connected to the cost regression it caused three jobs downstream.
LineageData platform cost joined to the cloud, Kubernetes and AI spend serving the same product, so cost per customer includes the warehouse credits, not just the compute.
ValueEach scenario carries severity, saving and evidence, and becomes a pull request against the repository that owns the resource, applied only after human approval.
Auto-suspend thresholds set generously and never revisited, warehouses sized for a quarterly peak and running all year, all-purpose clusters used for scheduled jobs. The cost of a warehouse idling is invisible on an invoice and obvious in system tables.
A single unpartitioned scan repeated hourly can outweigh an entire team's analytics. Cost by query, by user and by role makes the outlier attributable to someone who can fix it.
Run cost tracked per pipeline over time, so a schema change or data-volume shift that doubles the cost of a nightly job surfaces the week it happens rather than at contract renewal.
Time-travel retention, fail-safe storage, unused clones, and staging tables that outlived their migration. Storage is cheap per gigabyte and expensive by accumulation.
DBCU and capacity contract consumption paced against remaining term, so an under-consumed commitment is a negotiation input rather than a year-end write-off.
Generic warehouse rightsizing scratches the surface. The rows below are specific to how each platform actually bills and behaves.
| Platform | Billing grain ingested | Where the waste sits | AI workloads separated |
|---|---|---|---|
| Databricks | DBUs by job, cluster, warehouse and SKU; DBCU commitment drawdown | All-purpose clusters running scheduled jobs, generous auto-termination, photon-eligible workloads left off, small-file and compaction overhead | Mosaic AI serving and training |
| Snowflake | Credits by warehouse, query, user and role; capacity contract drawdown | Warehouses sized for a quarterly peak, unpartitioned repeated scans, time-travel and fail-safe retention, idle auto-suspend windows | Cortex and Vector Search |
| Google BigQuery | Slot-seconds and on-demand bytes by project, dataset and job; reservation utilization | On-demand pricing where flat-rate would be cheaper, missing partition pruning, full-table scans behind dashboards, unused slot reservations | Vertex-integrated workloads |
| MongoDB Atlas | Cluster tier, storage, backup and data transfer by project and organization | Tiers raised for an incident and never lowered, over-provisioned IOPS, backup retention beyond policy, cross-region transfer | Atlas Vector Search |
ClickHouse and Oracle are on the connector roadmap and will land at the same grain: the ingestion model is the same regardless of platform.
A point-in-time tuning exercise is right on the day it runs and quietly wrong three months later. DigiUsher re-evaluates continuously and tracks whether each saving is still being realized in the bill, so the number you reported in January is still true in March, or you know precisely when it stopped being true.
Illustrative shape. The last bar is the one most tools never report, and the gap between it and the first is why data platform costs return to baseline after a successful optimization project.
Scenarios run continuously against current usage rather than during a quarterly review, so a regression is caught in the week it appears.
Every change becomes a pull request against the repository that owns the resource. Your engineers approve; your pipeline applies. DigiUsher never holds write credentials and never rewrites a production query on its own authority.
Savings move identified → applied → verified-realized, then stay monitored. A saving that decays is reported as decayed rather than left on the running total.
A pipeline whose run cost doubled after a schema change is caught the week it happens, not at renewal. Because every domain shares one schema, this metric composes with the others into an AI-inclusive cost per customer: one formula, one auditable lineage.
Cost per customer can include this domain's cost alongside cloud, AI tokens, data credits, Kubernetes pods and SaaS seats. Point tools each compute a fraction; one schema computes the whole number.
Every optimization moves identified → applied → verified-realized, where the third state means subsequent billing data confirms the reduction. Most tools report the first and let you assume the third.
The full data model is exposed over the Model Context Protocol, so Claude, Copilot, Gemini or an in-house LLM can answer questions about this domain under the same role-based access control as the dashboards.
From system tables rather than invoice summaries: Databricks system tables for DBU consumption by cluster, job and SKU, and Snowflake ACCOUNT_USAGE and ORGANIZATION_USAGE for credits by warehouse, query, user and role. That grain is what makes per-pipeline and per-query attribution possible.
The total data platform cost attributable to one execution of a scheduled pipeline. Tracked over time it turns cost regression into a detectable event: a schema change or volume shift that doubles a nightly job is visible the week it happens.
Yes. Mosaic AI, Snowflake Cortex and Vector Search are treated as their own AI cost categories rather than absorbed into a general platform total, so AI spend inside the data platform is visible alongside AI spend everywhere else in the estate.
Databricks, Snowflake, MongoDB Atlas and Google BigQuery are available today, with ClickHouse and Oracle on the connector roadmap. All four land at system-table grain in the same FOCUS schema, so a DBU, a credit and a slot-second are directly comparable.
An invoice tells you what you spent. Query history tells you when. Neither explains why a job that cost $40 last month now costs $120. Ingesting at job, query, warehouse and dataset grain puts the cause next to the cost, and joins it to the pipeline and team that own it.
No. Recommendations, including query and configuration changes, become pull requests against the repository that owns the resource, reviewed by the engineers who own the pipeline and applied by your own CI/CD. DigiUsher never holds write credentials against a production data platform.
Because schemas change, code drifts and new teams spin up workloads, so a point-in-time tuning exercise stops being correct. DigiUsher re-evaluates continuously and monitors whether each verified saving is still holding in the bill, reporting decay rather than leaving it on the running total.
Yes. DBCU and capacity contract consumption is paced against the remaining term, so an under-consumed commitment becomes a negotiation input rather than a year-end surprise.
44 more answers: TVR, FinOps, AI cost, Kubernetes, allocation and vendor selection →
First insight within 48 hours of connecting a source.