Data Software Engineer – Spark, Python, Databricks (L2 / L4)
$21,000–$42,000 year
HybridBengaluru, Karnataka, India
Job Summary
Design and build distributed data processing systems using Spark and Hadoop, optimizing applications for performance and scalability. Create and manage ETL/ELT pipelines for large-scale data ingestion while developing real-time streaming systems with Spark Streaming or Storm and Kafka. Develop and optimize workloads on AWS or Azure Databricks, performing cluster management, job scheduling, and performance tuning. Integrate data from diverse sources including RDBMS, ERP, and file systems using Hive, Impala, and NoSQL stores like HBase and Cassandra. Write advanced Python scripts for data transformations and automations, leveraging strong SQL skills for validation and query optimization. Provide technical leadership and mentoring to junior engineers if joining as an L4 candidate. Collaborate within Agile teams on end-to-end solution design for Big Data platforms.
Required Qualifications
- Apache Spark – Expert level (core, SQL, streaming)
- Python – Strong hands-on
- Distributed computing fundamentals
- Hadoop ecosystem: Hadoop v2, MapReduce, HDFS, Sqoop
- Streaming systems: Spark Streaming / Storm
- Messaging: Kafka or RabbitMQ
- SQL – Advanced (joins, stored procedures, query optimization)
- NoSQL: HBase, Cassandra, MongoDB
- ETL frameworks & data pipeline design
- Hive / Impala querying
- Performance tuning of Spark jobs
- AWS or Azure Databricks
- Experience working in Agile
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.