Skip to content
Data platform economics

Domain edition

Data platform cost, at the grain where the answer lives.

DigiUsher ingests data platform cost from system tables rather than invoice summaries: DBUs by job and cluster, credits by query and warehouse, slot-seconds by job, Atlas tiers by project.

That grain is what turns a platform total into an attributable cost, and what lets a schema change upstream be connected to the cost regression it caused three jobs downstream.

4
Platforms available
2 more on the roadmap
Query
Ingestion grain
query · job · cluster · dataset
€1M
Realized in 45 days
European enterprise
Own
AI cost categories
Mosaic AI · Cortex · Vector Search

Available today system-table grain

  • Databricks Databricks system tables · DBU
  • Snowflake Snowflake ACCOUNT_USAGE · credits
  • MongoDB MongoDB Atlas tier · storage · transfer
  • Google BigQuery Google BigQuery slots · on-demand bytes

Coming soon on the connector roadmap

  • ClickHouse (coming soon)
  • Oracle Oracle (coming soon)

Every platform lands in the same FOCUS schema, so a Databricks DBU, a Snowflake credit and a BigQuery slot-second are directly comparable, and join to the cloud, Kubernetes and AI cost that surrounds them. A new platform is a connector, not a platform release.

Why grain matters

An invoice tells you what. System tables tell you when. Neither tells you why.

A monthly platform total cannot explain why Tuesday’s job costs three times what it cost last month. DigiUsher ingests at the grain where the answer lives, the job, the query, the warehouse, the cluster, and then joins that grain to the pipeline, dataset, team and downstream consumer it belongs to.

The cause of a cost spike is rarely where it surfaces. Attribution without lineage just relocates the argument.

  1. 01

    Query and job level

    DBUs by job and cluster, credits by query, user and role, slot-seconds by job. The single unpartitioned scan repeated hourly becomes attributable to someone who can fix it.

    Grain
  2. 02

    Compute level

    Warehouse and cluster configuration against observed utilization: auto-suspend thresholds, sizing, all-purpose clusters used for scheduled work, reservation and slot commitments.

    Grain
  3. 03

    Storage and dataset level

    Time-travel retention, fail-safe storage, unused clones, staging tables that outlived their migration, and Atlas tiers that were raised once and never lowered.

    Grain
  4. 04

    Pipeline and lineage level

    Cost rolled up per pipeline run and per dataset, so a schema change upstream is connected to the cost regression it caused three jobs downstream.

    Lineage
  5. 05

    Business level

    Data platform cost joined to the cloud, Kubernetes and AI spend serving the same product, so cost per customer includes the warehouse credits, not just the compute.

    Value
Waste scenarios

Where the money actually goes.

Each scenario carries severity, saving and evidence, and becomes a pull request against the repository that owns the resource, applied only after human approval.

Pipeline run cost before and after a schema change A line chart of nightly pipeline run cost over three weeks. Cost is flat at around four dollars per run, then steps up sharply after a schema change on day eleven to around nine dollars, and stays there. An invoice-level view shows only a small monthly rise; the run-level view shows the exact night the regression began. $0 $5 $10 NIGHTLY RUNS · 21 DAYS schema change $4.12 avg $9.34 avg · +127% Invoice view for the same period: platform spend up 9% month on month, indistinguishable from growth.
Run-level cost turns a regression into a dated event with an owner. At invoice grain the same change is a rounding difference nobody investigates until renewal.
  1. 01

    Cluster and warehouse discipline

    Auto-suspend thresholds set generously and never revisited, warehouses sized for a quarterly peak and running all year, all-purpose clusters used for scheduled jobs. The cost of a warehouse idling is invisible on an invoice and obvious in system tables.

  2. 02

    Query-level outliers

    A single unpartitioned scan repeated hourly can outweigh an entire team's analytics. Cost by query, by user and by role makes the outlier attributable to someone who can fix it.

  3. 03

    Pipeline regression detection

    Run cost tracked per pipeline over time, so a schema change or data-volume shift that doubles the cost of a nightly job surfaces the week it happens rather than at contract renewal.

  4. 04

    Storage lifecycle

    Time-travel retention, fail-safe storage, unused clones, and staging tables that outlived their migration. Storage is cheap per gigabyte and expensive by accumulation.

  5. 05

    Commitment drawdown

    DBCU and capacity contract consumption paced against remaining term, so an under-consumed commitment is a negotiation input rather than a year-end write-off.

Platform by platform

Each platform wastes money in its own dialect.

Generic warehouse rightsizing scratches the surface. The rows below are specific to how each platform actually bills and behaves.

Platform Billing grain ingested Where the waste sits AI workloads separated
Databricks DBUs by job, cluster, warehouse and SKU; DBCU commitment drawdown All-purpose clusters running scheduled jobs, generous auto-termination, photon-eligible workloads left off, small-file and compaction overhead Mosaic AI serving and training
Snowflake Credits by warehouse, query, user and role; capacity contract drawdown Warehouses sized for a quarterly peak, unpartitioned repeated scans, time-travel and fail-safe retention, idle auto-suspend windows Cortex and Vector Search
Google BigQuery Slot-seconds and on-demand bytes by project, dataset and job; reservation utilization On-demand pricing where flat-rate would be cheaper, missing partition pruning, full-table scans behind dashboards, unused slot reservations Vertex-integrated workloads
MongoDB Atlas Cluster tier, storage, backup and data transfer by project and organization Tiers raised for an incident and never lowered, over-provisioned IOPS, backup retention beyond policy, cross-region transfer Atlas Vector Search

ClickHouse and Oracle are on the connector roadmap and will land at the same grain: the ingestion model is the same regardless of platform.

The problem with a one-off tune

Optimization expires. Schemas change, code drifts, new teams spin up workloads.

A point-in-time tuning exercise is right on the day it runs and quietly wrong three months later. DigiUsher re-evaluates continuously and tracks whether each saving is still being realized in the bill, so the number you reported in January is still true in March, or you know precisely when it stopped being true.

Identified
100%
Applied
72%
Verified in bill
59%
Still holding at 90 days
48%

Illustrative shape. The last bar is the one most tools never report, and the gap between it and the first is why data platform costs return to baseline after a successful optimization project.

Continuous

Re-evaluated, not re-audited

Scenarios run continuously against current usage rather than during a quarterly review, so a regression is caught in the week it appears.

Governed

Reviewed before it applies

Every change becomes a pull request against the repository that owns the resource. Your engineers approve; your pipeline applies. DigiUsher never holds write credentials and never rewrites a production query on its own authority.

Verified

Confirmed in the bill

Savings move identified → applied → verified-realized, then stay monitored. A saving that decays is reported as decayed rather than left on the running total.

Signature value metric

Cost per pipeline run

A pipeline whose run cost doubled after a schema change is caught the week it happens, not at renewal. Because every domain shares one schema, this metric composes with the others into an AI-inclusive cost per customer: one formula, one auditable lineage.

Compose

Across domains

Cost per customer can include this domain's cost alongside cloud, AI tokens, data credits, Kubernetes pods and SaaS seats. Point tools each compute a fraction; one schema computes the whole number.

Verify

Savings that landed

Every optimization moves identified → applied → verified-realized, where the third state means subsequent billing data confirms the reduction. Most tools report the first and let you assume the third.

Ask

Conversationally

The full data model is exposed over the Model Context Protocol, so Claude, Copilot, Gemini or an in-house LLM can answer questions about this domain under the same role-based access control as the dashboards.

Frequently asked questions

Direct answers on data platform economics.

How does DigiUsher ingest Databricks and Snowflake cost?

From system tables rather than invoice summaries: Databricks system tables for DBU consumption by cluster, job and SKU, and Snowflake ACCOUNT_USAGE and ORGANIZATION_USAGE for credits by warehouse, query, user and role. That grain is what makes per-pipeline and per-query attribution possible.

What is cost per pipeline run?

The total data platform cost attributable to one execution of a scheduled pipeline. Tracked over time it turns cost regression into a detectable event: a schema change or volume shift that doubles a nightly job is visible the week it happens.

Does DigiUsher separate AI workloads inside data platforms?

Yes. Mosaic AI, Snowflake Cortex and Vector Search are treated as their own AI cost categories rather than absorbed into a general platform total, so AI spend inside the data platform is visible alongside AI spend everywhere else in the estate.

Which data platforms does DigiUsher support?

Databricks, Snowflake, MongoDB Atlas and Google BigQuery are available today, with ClickHouse and Oracle on the connector roadmap. All four land at system-table grain in the same FOCUS schema, so a DBU, a credit and a slot-second are directly comparable.

Why is system-table ingestion better than reading the invoice?

An invoice tells you what you spent. Query history tells you when. Neither explains why a job that cost $40 last month now costs $120. Ingesting at job, query, warehouse and dataset grain puts the cause next to the cost, and joins it to the pipeline and team that own it.

Does DigiUsher rewrite queries automatically?

No. Recommendations, including query and configuration changes, become pull requests against the repository that owns the resource, reviewed by the engineers who own the pipeline and applied by your own CI/CD. DigiUsher never holds write credentials against a production data platform.

Why do data platform savings disappear after a few months?

Because schemas change, code drifts and new teams spin up workloads, so a point-in-time tuning exercise stops being correct. DigiUsher re-evaluates continuously and monitors whether each verified saving is still holding in the bill, reporting decay rather than leaving it on the running total.

Can we track Databricks or Snowflake commitment drawdown?

Yes. DBCU and capacity contract consumption is paced against the remaining term, so an under-consumed commitment becomes a negotiation input rather than a year-end surprise.

44 more answers: TVR, FinOps, AI cost, Kubernetes, allocation and vendor selection →

Run it against your own data platform.

First insight within 48 hours of connecting a source.