About the Role
7+ years in data engineering with a deep Azure focus.
End-to-end ownership of Databricks and lakehouse architecture.
Proven ETL modernization at enterprise scale.
Native fluency across Azure Data Factory, Microsoft Fabric, Synapse, and ADLS Gen2.
Production-level programming in PySpark, Scala, and Python.
Key Responsibilities
ETL Modernization and Migration
Lead the end-to-end migration of ETL pipelines to Azure-native solutions.
Design and implement scalable data pipelines using Azure Data Factory and Databricks.
Architect and deliver a medallion lakehouse (Bronze, Silver, Gold) on Azure Data Lake Storage Gen2.
Ensure zero-downtime cutover strategies and backward compatibility during migration.
Azure Data Platform Engineering
Build and maintain production data pipelines across ADF, Synapse Analytics, and Databricks / Azure Fabric.
Implement Delta Lake for ACID transactions, schema enforcement, and time-travel capabilities.
Optimize Spark workloads for performance, cost, and reliability at scale.
Integrate streaming pipelines using Azure Event Hubs and Spark Structured Streaming.
Architecture and Design
Own the data architecture from ingestion to consumption, applying medallion and modern data warehouse patterns.
Evaluate and select the right Azure services for each layer of the data platform.
Collaborate with analysts and stakeholders to align platform design with business needs.
Drive decisions on partitioning strategies, compaction, Z-ordering, and query optimization in Delta Lake.
Governance, Security and DevOps
Implement Unity Catalog for data governance, lineage tracking, and access control.
Apply Azure security best practices: RBAC, Key Vault integration, and data encryption at rest and in transit.
Champion CI/CD using Git-based workflows and Infrastructure as Code (Terraform / Bicep).
Requirements
Core Azure Data Engineering
Azure Data Factory: pipeline authoring, parameterization, triggers, and integration runtimes.
Azure Databricks: cluster management, notebook development, job orchestration, and Delta Live Tables.
Azure Data Lake Storage Gen2: hierarchical namespace, access tiers, and lifecycle policies.
Azure Synapse Analytics: dedicated and serverless SQL pools, pipelines, and integration with Databricks.
Delta Lake: ACID transactions, schema evolution, merge operations, and performance tuning.
Microsoft Fabric: unified analytics across Lakehouse, Data Factory, and Real-Time Intelligence workloads.
Programming and Processing
Advanced PySpark and/or Scala for large-scale distributed data processing.
Python for scripting, automation, and data transformation logic.
SQL proficiency for complex analytics and data modelling.
Strong grasp of distributed computing, shuffle optimization, and caching strategies.
Architecture Patterns
Lakehouse architecture and medallion design (Bronze, Silver, Gold).
Modern data warehouse patterns and dimensional modelling.
Event-driven architecture and real-time streaming pipelines.
DevOps and Collaboration
CI/CD for data pipelines: GitHub Actions, Azure DevOps, or equivalent.
Infrastructure as Code: Terraform or Azure Bicep.
Orchestration: ADF, Apache Airflow, or Databricks Workflows.
Strong Git practices: branching strategies, PR reviews, and collaborative development.