# 𝐄𝐧𝐝-𝐭𝐨-𝐄𝐧𝐝 𝐃𝐚𝐭𝐚 𝐄𝐧𝐠𝐢𝐧𝐞𝐞𝐫𝐢𝐧𝐠 𝐏𝐢𝐩𝐞𝐥𝐢𝐧𝐞 𝐨𝐧 𝐃𝐚𝐭𝐚𝐛𝐫𝐢𝐜𝐤𝐬 (𝐅𝐌𝐂𝐆 𝐃𝐨𝐦𝐚𝐢𝐧)!

𝐉𝐮𝐬𝐭 𝐛𝐮𝐢𝐥𝐭 𝐚𝐧 𝐄𝐧𝐝-𝐭𝐨-𝐄𝐧𝐝 𝐃𝐚𝐭𝐚 𝐄𝐧𝐠𝐢𝐧𝐞𝐞𝐫𝐢𝐧𝐠 𝐏𝐢𝐩𝐞𝐥𝐢𝐧𝐞 𝐨𝐧 𝐃𝐚𝐭𝐚𝐛𝐫𝐢𝐜𝐤𝐬 (𝐅𝐌𝐂𝐆 𝐃𝐨𝐦𝐚𝐢𝐧)!

When two companies merge, unifying their messy, disparate data systems is always a major operational bottleneck. I recently tackled a real-world merger and acquisition scenario: integrating transactional and master data from an acquired startup (Sports Bar) into an established enterprise platform (Atlon).

Here is an overview of the technical architecture and pipeline workflow I designed:

𝐀𝐫𝐜𝐡𝐢𝐭𝐞𝐜𝐭𝐮𝐫𝐞 & 𝐏𝐢𝐩𝐞𝐥𝐢𝐧𝐞 𝐎𝐯𝐞𝐫𝐯𝐢𝐞𝐰 (𝐌𝐞𝐝𝐚𝐥𝐥𝐢𝐨𝐧 𝐀𝐫𝐜𝐡𝐢𝐭𝐞𝐜𝐭𝐮𝐫𝐞)

𝐈𝐧𝐠𝐞𝐬𝐭𝐢𝐨𝐧 𝐋𝐚𝐲𝐞𝐫 (𝐀𝐖𝐒 𝐒𝟑 & 𝐔𝐧𝐢𝐭𝐲 𝐂𝐚𝐭𝐚𝐥𝐨𝐠): Ingested daily raw transaction batches (CSV) and dimension data into AWS S3 buckets, connected seamlessly to Databricks using external locations and Unity Catalog.

𝐁𝐫𝐨𝐧𝐳𝐞 𝐋𝐚𝐲𝐞𝐫 (𝐑𝐚𝐰 𝐈𝐧𝐠𝐞𝐬𝐭𝐢𝐨𝐧): Stored untransformed data in Delta format with metadata tracking (file\_name, read\_timestamp, file size) and Change Data Feed (CDF) enabled.

𝐒𝐢𝐥𝐯𝐞𝐫 𝐋𝐚𝐲𝐞𝐫 (𝐂𝐥𝐞𝐚𝐧𝐢𝐧𝐠 & 𝐄𝐧𝐫𝐢𝐜𝐡𝐦𝐞𝐧𝐭):

Handled data deduplication, null-value imputations, and text casing standardization.

Standardized city typos and parsed unstructured text to extract product variants using Regex in PySpark.

Generated deterministic surrogate keys via SHA hashing to resolve product identifier discrepancies across systems. This is a new thing I learned: (𝐒𝐇𝐀 𝐡𝐚𝐬𝐡𝐢𝐧𝐠 is a mathematical algorithm that takes any input data (like text or numbers) and converts it into a unique, fixed-length string of characters)

𝐆𝐨𝐥𝐝 𝐋𝐚𝐲𝐞𝐫 (𝐁𝐮𝐬𝐢𝐧𝐞𝐬𝐬 𝐑𝐞𝐚𝐝𝐲 & 𝐂𝐨𝐧𝐟𝐨𝐫𝐦𝐞𝐝 𝐌𝐚𝐫𝐭𝐬):

Implemented PySpark Window functions to derive latest gross pricing. Aggregated daily transaction granularity to monthly level to match enterprise reporting standards. Executed incremental Upsert (MERGE) operations to reconcile historical backfills and daily incoming batches with the parent dataset.

𝐎𝐫𝐜𝐡𝐞𝐬𝐭𝐫𝐚𝐭𝐢𝐨𝐧: Automated the entire sequence (Dimensions ➔ Facts ➔ Staging Cleanup) using Databricks Workflows/Jobs

𝘼𝙣𝙖𝙡𝙮𝙩𝙞𝙘𝙨 & 𝘼𝙄 𝙎𝙚𝙧𝙫𝙞𝙣𝙜: 𝙔𝙚𝙩 𝙩𝙤 𝙞𝙢𝙥𝙡𝙚𝙢𝙚𝙣𝙩

𝐊𝐞𝐲 𝐓𝐚𝐤𝐞𝐚𝐰𝐚𝐲𝐬 & 𝐈𝐦𝐩𝐚𝐜𝐭

Delivered a single source of truth for unified executive revenue & product KPI tracking across merged entities.

Handled schema evolution, idempotency, and automated daily incremental updates with zero manual intervention.

A huge shoutout to codebasics for the fantastic practical blueprint!

#DataEngineering #Databricks #PySpark #DeltaLake #AWS #MedallionArchitecture #DataPipeline #ETL
