Skip to content

Case Studies Technology & Cloud Services

Re-engineering legacy SQL Server ETL into a Databricks-native pipeline

A global technology and cloud services company needed to move off an on-prem SQL Server ETL estate that could no longer handle its data volumes or processing SLAs.

~400GB processed dailyFull pipeline run in under 2 hoursMaterially higher throughput than the legacy estate

Client challenge

What was wrong with the legacy environment

On-prem SQL Server ETL pipelines could not scale to growing data volumes

Processing windows regularly breached agreed SLAs

High infrastructure overhead to keep the legacy estate running

Limited ability to parallelise or scale processing on demand

Modernization approach

How Karsient approached the transformation

Karsient re-engineered the SQL-based ETL logic into Databricks-native PySpark and Scala workloads, replacing the legacy processing engine while preserving business logic and output parity.

Architecture

The modern target architecture

1

A hub-and-spoke ingestion pattern feeding a central Lakehouse

2

ELT pipelines built on Databricks replacing SQL Server stored-procedure ETL

3

Parallelised, distributed processing replacing single-node batch jobs

4

Automated validation comparing legacy and new pipeline outputs

Migration

Migration & re-engineering strategy

SQL ETL logic was converted and validated table by table, run alongside the legacy pipeline until outputs matched exactly, before the on-prem jobs were decommissioned.

Engineering improvements

What changed under the hood

SQL stored procedures converted to distributed PySpark/Scala jobs

Parallel processing replacing single-threaded legacy execution

Automated data validation between legacy and modernised outputs

Reduced infrastructure footprint versus the on-prem estate

Business impact

Measurable outcomes

~400GB of data processed reliably on a daily basis

Full processing pipeline completing in under 2 hours

Ability to handle substantially larger data volumes than before

Lower infrastructure overhead versus the legacy on-prem platform

Technology

Technology used

Databricks
Apache Spark
Python
SQL
Delta Lake

Want results like these?

Let's scope your next data or AI initiative