Skip to content
VERACITY
Energy utility

Smart-meter telemetry lakehouse on Scala Spark

Billions of half-hourly meter reads processed daily with a Scala Spark pipeline on Databricks, the foundation for settlement, network planning and forecasting.

2024·8 months·4 engineers, 1 tech lead
  • ~2.4 billion

    Records/day

  • 99.6%

    SLA met

  • 80+

    Legacy jobs retired

The challenge

The existing pipeline was a decade-old mix of SSIS packages and shell scripts. It missed its SLA more often than it met it, and every new use case, settlement, network planning, load forecasting, added another downstream extract. The data science team wanted feature stores; the ops team wanted reliability. Neither was getting either.

What we did

  • Rebuilt the ingestion tier as a Scala Spark job on Databricks, with strongly-typed Datasets and property-based tests via ScalaCheck.
  • Landed data into Delta Lake with medallion layering and time-travel snapshots for reproducible analytics.
  • Introduced Unity Catalog for lineage and fine-grained access, one place to answer 'who can see what'.
  • Deployed with Databricks Asset Bundles, environment-promoted through Azure DevOps with reproducible clusters.

Results

  • Nightly SLA now met 99.6% of the time (previously ~82%).
  • Downstream teams work from the lakehouse, the SSIS/shell estate is gone.
  • Data science team gets time-travel snapshots for reproducible model training.
  • Onboarding a new use case has dropped from months to weeks.

This is the first pipeline we've had that we're not afraid to change. That sounds small, it isn't.

Head of Data Platform · UK Energy Utility
Let's talk

Ready to build?

Tell us what you need built. We'll reply within one working day with an honest answer.