Staff Data Engineer (8-10 years' exp - Java/Python, Scala, Spark, Hadoop)
HybridBengaluru, Karnataka, India
Job Summary
Design and implement scalable distributed data pipelines using Spark, Kafka, and Delta Lake to bridge legacy SQL Server ecosystems with modern Hadoop and Databricks platforms. Optimize existing SQL Server models and define long-term migration strategies aligned with data governance standards while partnering with Agentic AI teams to provision LLM data and RAG pipelines. Provide guidance on database performance tuning, replication, and high-availability solutions for junior engineers, ensuring compliance with Visa's data privacy and encryption frameworks. Lead proof-of-concept initiatives to evaluate new data engineering technologies and collaborate cross-functionally on architecture decisions for AI-driven insights.
Required Qualifications
- Bachelor's or Master's degree in Computer Science, Data Engineering, or related field
- 8+ years of hands-on data engineering experience, including enterprise-scale data warehousing and pipeline design
- Proven success supporting and modernizing SQL Server–based data warehouses in high-SLA environments
- Production-level experience architecting Hadoop, Spark, and Databricks data pipelines
- Expertise in ETL/ELT frameworks, data modeling, and schema design for analytical and operational use cases
- Strong programming proficiency in Python, Java, or Scala
- Hands-on experience with AWS or Azure (Glue, Synapse, Redshift, Delta Lake, S3)
- Familiarity with Kafka, Airflow, Kubernetes, and containerized data services
- Understanding of RAG pipelines, vector databases, and AI data flows
- Experience designing data pipelines and APIs compatible with Model Context Protocol (MCP)-based agent frameworks, enabling seamless integration between AI agents, data services, and enterprise APIs
- Strong SQL optimization, debugging, and production troubleshooting experience
- Visa requires at least 3 days in office
Desired Qualifications
- Experience developing data warehouse migration or modernization from SQL Server to Hadoop/Spark ecosystems
- Deep understanding of data lineage, metadata management, and governance frameworks (e.g., Atlas, Great Expectations)
- Familiarity with LangChain, LangGraph, and MCP for integrating AI agents with data systems
- Strong ability to balance innovation and stability across coexisting legacy and modern data architectures
- Proven track record mentoring engineers and collaborating across global teams
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.