Iterative
Senior Data Engineer (Azure)
On-sitePosted 2 mo. ago
About the role
About the Role
We are looking for a Senior Data Engineer to design, build, and operate large-scale data pipelines and analytical solutions on the Microsoft Azure stack. You will own end-to-end ETL/ELT workflows that move data from operational systems into curated data lakes and warehouses, enabling analytics, reporting, and ML use cases across the business. The ideal candidate combines strong data modeling instincts with hands-on Azure platform expertise and a bias for reliable, observable pipelines.
Responsibilities
- Design and implement batch and incremental ETL/ELT pipelines using Azure Data Factory, Synapse Pipelines, and/or Azure Databricks (PySpark).
- Build and maintain data lakehouse architectures on Azure Data Lake Storage Gen2 using a medallion (bronze/silver/gold) pattern.
- Develop scalable transformations in SQL, Python, and PySpark; enforce schema, data quality, and lineage standards.
- Model and load dimensional and star-schema datasets into Synapse Dedicated SQL Pool, Azure SQL, or Databricks Delta tables for downstream analytics.
- Orchestrate complex dependencies, parameterized pipelines, event-driven triggers, and reusable frameworks with proper logging, alerting, and retry semantics.
- Monitor pipeline health, performance, and cost; tune Spark clusters, SQL pool DWUs, and ADF integration runtimes.
- Partner with analytics, BI, and platform teams to translate business requirements into reliable, testable data products.
- Implement CI/CD for data assets using Git, Azure DevOps/GitHub Actions, and Infrastructure-as-Code (ARM/Bicep/Terraform).
Requirements
- 5+ years of data engineering experience with a strong ETL/ELT foundation.
- Hands-on expertise across the Azure data stack: Azure Data Factory, Azure Synapse Analytics, Azure Data Lake Storage Gen2, Azure SQL, and Azure Databricks.
- Strong SQL skills including window functions, CTEs, query performance tuning, and dimensional modeling (Kimball/star schema).
- Proficiency in Python and/or PySpark for distributed data transformation.
- Practical experience with incremental load patterns: Change Data Capture (CDC), watermarking, MERGE/UPSERT logic, and handling late-arriving data.
- Solid understanding of file formats (Parquet, Delta, Avro, JSON), partitioning strategies, and compression trade-offs.
- Familiarity with orchestration patterns, parameterization, and error handling at scale.
- Experience integrating data from relational sources (SQL Server, PostgreSQL, MySQL), REST APIs, and event streams.
Nice to Have
- Azure certifications such as DP-203 (Azure Data Engineer Associate) or DP-300.
- Experience with Azure Purview/Microsoft Fabric, Unity Catalog, or modern data governance practices.
- Exposure to streaming ingestion (Event Hubs, Kafka, Azure Stream Analytics).
- Infrastructure-as-Code with Bicep or Terraform for Azure data platforms.
- Familiarity with dbt, Great Expectations, or other data testing/quality frameworks.
- Experience supporting Power BI or downstream semantic models.
Interview Questions
- Walk me through how you would design an end-to-end ETL pipeline that ingests 50 million rows nightly from an on-premises SQL Server source into a Synapse Dedicated SQL Pool using Azure Data Factory and Data Lake Storage Gen2. What format, partitioning, and load pattern would you choose, and why?
- Compare Azure Synapse Dedicated SQL Pool and Azure Databricks for a large-scale transformation workload. When would you pick one over the other, and how do you decide where the transformation logic lives in a medallion architecture?
- Describe how you implement incremental loads when the source system does not support CDC. How do you handle late-arriving rows, deletes, and pipeline restarts mid-run?
- You notice a nightly ADF pipeline that used to finish in 45 minutes is now taking 3 hours and occasionally timing out. Walk me through your troubleshooting methodology — what telemetry do you inspect first, and what are the most common root causes you would expect?
- Schema drift is one of the most common failure modes in production data pipelines. How do you detect, handle, and protect downstream consumers from breaking when a source system adds, renames, or removes columns?
- PySpark code is correct in dev but in production a stage fails with an out-of-memory error on the driver and skewed shuffle on executors. How would you diagnose and remediate this without doubling infrastructure cost?
- Beyond performance, how do you measure the reliability and trustworthiness of your data pipelines? Describe the data quality checks, observability signals, and alerting you would put in place for a critical daily mart.
- Our Azure spend on the data platform has grown 40% quarter-over-quarter. Where would you look first to cut cost without sacrificing SLAs, and which levers on ADF, Synapse, and Databricks typically have the biggest impact?