Energy utility
Smart-meter telemetry lakehouse on Scala Spark
Billions of half-hourly meter reads processed daily with a Scala Spark pipeline on Databricks, the foundation for settlement, network planning and forecasting.
2024·8 months·4 engineers, 1 tech lead
~2.4 billion
Records/day
99.6%
SLA met
80+
Legacy jobs retired
The challenge
The existing pipeline was a decade-old mix of SSIS packages and shell scripts. It missed its SLA more often than it met it, and every new use case, settlement, network planning, load forecasting, added another downstream extract. The data science team wanted feature stores; the ops team wanted reliability. Neither was getting either.
What we did
- Rebuilt the ingestion tier as a Scala Spark job on Databricks, with strongly-typed Datasets and property-based tests via ScalaCheck.
- Landed data into Delta Lake with medallion layering and time-travel snapshots for reproducible analytics.
- Introduced Unity Catalog for lineage and fine-grained access, one place to answer 'who can see what'.
- Deployed with Databricks Asset Bundles, environment-promoted through Azure DevOps with reproducible clusters.
Results
- Nightly SLA now met 99.6% of the time (previously ~82%).
- Downstream teams work from the lakehouse, the SSIS/shell estate is gone.
- Data science team gets time-travel snapshots for reproducible model training.
- Onboarding a new use case has dropped from months to weeks.
“This is the first pipeline we've had that we're not afraid to change. That sounds small, it isn't.”
Let's talk
Ready to build?
Tell us what you need built. We'll reply within one working day with an honest answer.