Anrgi Tech logo
Anrgi TechPosted 2 weeks ago

Senior Data Engineer (APAC Region)

$1,600,000–$2,000,000 year

On-sitePune, Maharashtra, India

Full TimeSenior LevelStartup

Job Summary

Design, build, and operate scalable data pipelines on Databricks from ingestion through transformation to consumption. Refactor legacy code and ETL processes to PySpark and modern ELT patterns while ensuring backward compatibility and data integrity. Implement robust error handling, monitoring, and alerting mechanisms to maintain pipeline reliability and performance. Develop reusable code libraries, enforce software engineering best practices, and troubleshoot production issues. Collaborate with data architects, analysts, Infra, Apps, and Cyber teams to align on requirements and share knowledge. Mentor junior engineers and document technical solutions and migration playbooks. Requires 5+ years in data engineering, 2-3 years with Databricks, mandatory Databricks certification, and proficiency in PySpark, Delta Lake, and Python.

Required Qualifications

  • Minimum 5 years in data engineering or related roles
  • At least 2-3 years of hands-on experience with Databricks platform
  • Proven track record of refactoring legacy code to modern frameworks
  • Experience building and maintaining production data pipelines at scale
  • Background working across multiple data sources and formats
  • Experience in agile development environments
  • Required Certifications - mandatory to have at least one certification
  • Databricks Certified Data Engineer Associate OR Databricks Certified Data Engineer Professional

Desired Qualifications

  • Strong foundation in data engineering principles, ETL/ELT processes, and data pipeline design patterns
  • Proven hands-on experience developing data pipelines using PySpark, including DataFrames API, Spark SQL, and performance optimization
  • Practical experience with Databricks workspace, cluster management, notebooks, and job orchestration
  • Knowledge of Databricks Workspace AI Agent capabilities and integration
  • Experience implementing data models including dimensional modeling, data vault, or lakehouse architectures
  • Understanding of Delta Lake features including ACID transactions, schema evolution, and optimization techniques
  • Strong Python programming skills for data processing and automation
  • SQL proficiency for data querying and transformation
  • Experience with cloud platforms (Azure, AWS, or GCP)
  • Understanding of data governance and security best practices
  • Knowledge of streaming data processing (Structured Streaming)
  • Familiarity with DevOps practices and CI/CD pipelines
  • Experience with version control systems (Git)
  • Understanding of data quality frameworks and testing methodologies
  • Databricks Certified Associate Developer for Apache Spark
  • Cloud platform certifications (Azure Data Engineer Associate, AWS Certified Data Analytics, or Google Cloud Professional Data Engineer)
  • Relevant data engineering or big data certifications
  • Strong problem-solving and analytical thinking abilities
  • Excellent communication skills to explain technical concepts clearly
  • Ability to work collaboratively in cross-functional teams
  • Self-motivated with strong attention to detail
  • Adaptable to changing priorities and technologies
  • Client-focused mindset with commitment to quality delivery

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce