Senior Developer, Data Engineer
On-siteHyderabad, Telangana, India
Job Summary
Own Apache Airflow end-to-end by writing DAGs that handle complex multi-step workflows across ingestion, transformation, validation, and delivery, ensuring idempotency, backfill safety, and SLA alerting. Manage Airflow deployments on Kubernetes, tune resource limits, and resolve scheduler performance issues before they become incidents. Build fault-tolerant ETL and ELT pipelines from diverse sources like APIs, message queues, and object storage, embedding schema validation, row-count reconciliation, and anomaly detection from the start. Design and maintain a lakehouse architecture using Apache Iceberg with a Bronze, Silver, and Gold medallion structure, while utilizing Databricks for cluster compute, Delta Lake workflows, and Unity Catalog. Develop dbt projects alongside orchestration layers, creating well-documented models, tests, and freshness checks. Contribute reusable components such as pipeline templates and custom operators to the shared platform, and partner with data scientists to deliver feature pipelines with high reproducibility. Instrument work with metrics, dashboards, and alerts, owning runbooks and on-call coverage.
Required Qualifications
- Apache Airflow: Expert level, with at least three years running Airflow in production
- Apache Airflow: knowledge of the scheduler internals, executor types, XComs, Connections, Variables, pools, and the TaskFlow API
- Airflow on Kubernetes: Real experience deploying and operating Airflow on Kubernetes via KubernetesExecutor or CeleryKubernetesExecutor
- Airflow on Kubernetes: knowledge of pod templates, resource tuning, persistent volume claims, Helm chart management, and upgrades that do not take the scheduler offline
- DAG Development: Strong Python applied to DAG authoring: dynamic task mapping, custom operators and sensors, cross-DAG triggers, parameterised pipelines, and proper test coverage
- ETL and ELT Pipelines: Demonstrated experience building batch and incremental pipelines at scale
- ETL and ELT Pipelines: understanding of transformation logic, data quality validation, and lineage capture
- Data Lakehouse Architecture: A solid grasp of lakehouse design in practice: how to implement a medallion model, when to use which table format, and how to make data reliably available to downstream consumers
- Apache Iceberg: Hands-on knowledge of the Iceberg table spec: snapshot isolation, partition evolution, hidden partitioning, and time-travel queries
- Apache Iceberg: experience managing Iceberg tables in a real catalog environment, whether Hive metastore, a REST catalog, or otherwise
- Databricks: Practical Databricks experience covering cluster management, Delta Lake, Workflows or Jobs, and Unity Catalog
- dbt: Confident building dbt projects from the ground up: models, sources, tests, snapshots, seeds, and macros
- dbt: knowledge of how to wire dbt into Airflow, write documentation that people actually use, and make freshness checks meaningful
- Python: Strong engineering-grade Python: well-structured, tested, and packaged
- Python: Comfort with pandas, PySpark, SQLAlchemy, and Pydantic
Desired Qualifications
- Apache Kafka and Streaming: Experience with near-real-time ingestion using Kafka, Kafka Connect, or Flink, including landing data into lakehouse storage
- Apache Spark: PySpark proficiency for large-scale distributed processing, with a feel for memory management, shuffle optimisation, and partition sizing
- Kubernetes and OpenShift: Platform knowledge beyond Airflow: Helm, RBAC, and namespace management, enough to self-serve on infrastructure without leaning on platform teams
- Data Governance and Contracts: Familiarity with data contract approaches, catalogue tools such as DataHub or Collibra, and quality frameworks like Great Expectations or Soda
- Cloud Platforms: Experience with managed data services on AWS, Azure, or GCP: S3, Glue, ADLS, Synapse, GCS, Dataproc, and similar
- Infrastructure as Code: Ability to provision data infrastructure using Terraform or Pulumi
- ML Feature Engineering: An understanding of feature stores and experience building pipelines that serve model training and inference reliably
- Databricks: Experience with Databricks Asset Bundles or MLflow
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.