BookMyShow logo
BookMyShowPosted 3 weeks ago

Software Development Engineer - I (Data)

On-siteMumbai, Maharashtra, India

Full TimeAssociates DegreeEnterprise

Job Summary

Design and maintain scalable ETL/ELT pipelines on Databricks using Spark, Delta Lake, and Unity Catalog to operationalize machine learning models. Develop production feature pipelines ensuring reliability and freshness, while collaborating with Data Science teams to deploy batch and real-time models via Databricks Model Serving. Optimize Spark jobs and SQL warehouses for performance, and implement system-level observability for pipelines and ML jobs including usage, cost, and data quality monitoring. Own Unity Catalog governance for datasets and features, and contribute to architecture decisions around lakehouse design and streaming tradeoffs. Perform analysis on built data to validate metrics and spot quality issues, while evaluating LLM and agentic frameworks within governance constraints.

Required Qualifications

  • 1-3 years of data engineering experience working on Databricks in production
  • Strong proficiency in PySpark/Spark SQL and Python
  • Solid understanding of Delta Lake, Unity Catalog, Lakeflow/DLT, and Databricks system tables (billing, compute, query history)
  • Experience building and maintaining feature pipelines or ML data infrastructure (feature stores, training/serving data parity)
  • Strong SQL skills and experience with warehouse performance tuning (query optimization, materialization strategies, cluster sizing)
  • Understanding of ML fundamentals — enough to have real conversations with Data Scientists about features, drift, and model lifecycle (you don't need to be building models yourself, but you should understand what "good" looks like)

Desired Qualifications

  • Experience with real-time/streaming architectures (Structured Streaming, Kafka, Lakebase or similar)
  • Exposure to LLM/agentic tooling (Databricks Genie, RAG pipelines, vector search)
  • Experience with cost governance/FinOps for Databricks workloads
  • Background with experimentation platforms or A/B testing infrastructure
  • Familiarity with orchestration tools (Databricks Jobs, Airflow) and CI/CD for data pipelines

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce