Senior Data Engineer
On-siteHyderabad, Telangana, India
Job Summary
Lead the design, development, and delivery of scalable batch and real-time ETL/ELT pipelines, owning complex data solutions from requirements through production support. Build cloud-based lakehouse solutions using Databricks and AWS, integrating structured, semi-structured, and unstructured data from enterprise, manufacturing, API, and third-party sources. Optimize Spark workloads, implement workflow orchestration, monitoring, and data-quality controls, while developing CI/CD pipelines and automated testing. Define data models, metadata management, and governance frameworks, then lead architecture reviews, code reviews, and root-cause analysis. Mentor engineers and collaborate with architects, data scientists, and DevOps teams to support technical roadmaps and delivery-risk management. Participate in occasional off-hours operational support. Experience in biotechnology, pharmaceutical, or life sciences is preferred.
Required Qualifications
- Master's degree in Computer Science, Engineering, Information Technology, Data Science, or a related field and at least 7 years of relevant experience
- Bachelor's degree in a related field and at least 9 years of relevant experience
- Advanced hands-on experience with Databricks, Apache Spark, PySpark, Spark SQL, Delta Lake, Python, and SQL
- Experience designing and operating production-grade batch and streaming pipelines
- Strong understanding of distributed computing, lakehouse architecture, data warehousing, and data integration
- Experience with Databricks Workflows or comparable orchestration tools
- Strong experience with AWS data, compute, storage, security, and monitoring services
- Experience with Spark performance tuning, cluster optimization, partitioning, and cost management
- Experience with Git, CI/CD, automated testing, monitoring, and production deployment
- Experience implementing data quality, metadata management, lineage, governance, and access controls
- Strong understanding of RBAC, least privilege, encryption, auditability, and regulated-data requirements
- Experience leading technical design, code reviews, and complex production implementations
- Ability to define reusable patterns, engineering standards, and development best practices
- Strong communication, collaboration, mentoring, and problem-solving skills
- Experience working in Agile or Scaled Agile delivery environments
Desired Qualifications
- Experience in biotechnology, pharmaceutical, life sciences, manufacturing, or another regulated industry
- Experience with Unity Catalog, data products, Data Fabric, Data Mesh, or similar enterprise data architectures
- Experience building APIs and secure data services
- Experience with relational, NoSQL, analytical, operational, or vector databases
- Experience with OLAP and OLTP data modeling and performance tuning
- Experience with Kafka, Kinesis, or other streaming technologies
- Experience supporting AI and Generative AI solutions, including RAG, embeddings, vector search, and governed enterprise data access
- Familiarity with AI-assisted development tools such as GitHub Copilot, OpenAI Codex, or equivalent platforms
- AWS Certified Data Engineer or another relevant AWS certification
- Databricks Certified Data Engineer Associate or Professional
- Scaled Agile Framework certification
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.