Architect - Data Engineer
HybridHyderabad, Telangana, India
Job Summary
Design scalable and resilient data processing solutions using Apache Spark and modern cloud platforms, creating architecture designs and technical specifications for distributed data systems. Guide engineering teams on architecture decisions, conduct design reviews, and establish standards for observability, performance optimization, and operational excellence. Collaborate with product managers and technical leads to translate business requirements into scalable technical solutions while mentoring engineers on distributed systems design and Spark best practices. Work hybrid, reporting two days in-person at an assigned TU office location.
Required Qualifications
- 10+ years of experience in software engineering, data engineering, or distributed systems development
- Strong hands-on expertise with Apache Spark (Spark SQL, Structured Streaming, DataFrames, Dataset APIs)
- Experience designing and building large-scale distributed data processing systems
- Strong knowledge of Spark optimization techniques including: Partitioning Strategies Shuffle Optimization Join Optimization Memory Management Resource Utilization Performance Tuning
- Proficiency in Scala, Java, or Python
- Experience with technologies such as: Hadoop Hive Iceberg AWS EMR AWS Glue GCP Dataproc BigQuery
- Experience designing and operating batch and streaming data pipelines
- Understanding of cloud-native architecture and distributed systems principles
- Experience deploying Spark solutions on AWS and/or GCP platforms
- This is a hybrid position and involves regular performance of job responsibilities virtually as well as in-person at an assigned TU office location for a minimum of two days a week
Desired Qualifications
- Experience with AWS services such as EMR, Glue, S3, Lambda, EKS, ECS, Step Functions, and CloudWatch
- Experience with GCP services such as Dataproc, BigQuery, Cloud Storage, Dataflow, Pub/Sub, and Composer
- Experience managing terabyte-to-petabyte scale data processing environments
- Experience with Spark Structured Streaming, real-time analytics, and event-driven data processing
- Contributions to Spark optimization, platform engineering, or open-source data ecosystem projects are a plus
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.