Konovo logo
KonovoPosted 1 month ago

Data Engineer

HybridBengaluru, Karnataka, India

Full TimeSmall

Job Summary

Build and own production-grade batch ELT pipelines in Databricks/Spark running on a 15-minute cadence during working hours and hourly off-hours. Design, scale, and operate the lakehouse using medallion patterns, handling incremental loads, backfills, and late-arriving data. Create curated, well-documented data products trusted for decision-making while improving reliability through automated testing, monitoring, and clear SLAs. Partner with stakeholders to translate needs into durable data models and ensure governance, security, and lineage standards are met. Work hybrid in Bengaluru with overlapping hours for EST teams.

Required Qualifications

  • 5+ years building and operating production data pipelines (data engineering—not primarily BI/analytics or data science)
  • Strong SQL plus strong data modeling skills (dimensional and/or lakehouse modeling)
  • Hands-on Databricks + Spark experience in production (debugging, performance tuning, cost awareness)
  • Experience building batch ELT at frequent cadence (e.g., 15-minute schedules), including idempotency, backfills, and late-arriving data patterns
  • Experience with orchestration and transformation tooling (e.g., Airflow + dbt) and modern development practices (version control, CI/CD)
  • Strong data quality discipline: automated tests/expectations, monitoring/alerting, and clear SLAs for critical datasets
  • Ownership mindset: you build it, you run it (triage, incident response, continuous improvement)
  • Clear written and verbal communication for engineering work (requirements clarification, design docs, tradeoffs)
  • Comfortable working hybrid in Bengaluru (Whitefield) and overlapping 3–4 hours with EST

Desired Qualifications

  • Deep Delta Lake experience (schema evolution, OPTIMIZE/Z-ORDER, compaction, partitioning strategy)
  • CDC ingestion patterns (e.g., Debezium/Fivetran/HVR or custom CDC) and handling late/out-of-order events
  • Streaming or near-real-time pipelines (Structured Streaming, Kafka, Auto Loader) even if the core role is batch
  • Strong observability practices for data systems (metrics, lineage, data contracts, incident postmortems)
  • Cost and performance optimization in Databricks (cluster sizing, job tuning, Photon, caching strategies)
  • Experience building governed data products for multi-tenant consumption (RBAC, PII handling, auditability)
  • Exposure to healthcare, life sciences, or market research data and related compliance considerations

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce