Senior Fullstack Developer
On-siteSaidapet, State of Tamil Nādu, Republic of India
Job Summary
Design, build, and optimize scalable data pipelines using PySpark, Databricks, and SQL on AWS cloud platforms. Collaborate with analysts and scientists to implement batch and streaming ingestion frameworks, performing transformation, cleansing, and validation with Python. Develop reusable ETL/ELT components, maintain data models and marts, and ensure robust testing with proper logging and alerting. Optimize distributed processing workflows while applying best practices in version control, code reviews, and agile development. Leverage AWS services for lakehouse architecture and participate in data governance to ensure compliance with privacy and security standards. Contribute to documentation of processes, workflows, and metadata using data catalogs. Drive continuous improvement in engineering practices and automation to increase delivery quality within a hybrid, collaborative environment focused on life sciences AI solutions.
Required Qualifications
- 4 to 6 years of professional experience in Data Engineering or a related field
- Strong programming experience with Python and experience using Python for data wrangling, pipeline automation, and scripting
- Deep expertise in writing complex and optimized SQL queries on large-scale datasets
- Solid hands-on experience with PySpark and distributed data processing frameworks
- Expertise working with Databricks for developing and orchestrating data pipelines
- Experience with AWS cloud services such as S3, Glue, EMR, Athena, Redshift, and Lambda
- Practical understanding of ETL/ELT development patterns and data modeling principles (Star/Snowflake schemas)
- Experience with job orchestration tools like Airflow, Databricks Jobs, or AWS Step Functions
- Understanding of data lake, lakehouse, and data warehouse architectures
- Familiarity with DevOps and CI/CD tools for code deployment (e.g., Git, Jenkins, GitHub Actions)
- Strong troubleshooting and performance optimization skills in large-scale data processing environments
- Excellent communication and collaboration skills, with the ability to work in cross-functional agile teams
Desired Qualifications
- AWS or Databricks certifications (e.g., AWS Certified Data Analytics, Databricks Data Engineer Associate/Professional)
- Exposure to data observability, monitoring, and alerting frameworks (e.g., Monte Carlo, Datadog, CloudWatch)
- Experience working in healthcare, life sciences, finance, or another regulated industry
- Familiarity with data governance and compliance standards (GDPR, HIPAA, etc.)
- Knowledge of modern data architectures (Data Mesh, Data Fabric)
- Exposure to streaming data tools like Kafka, Kinesis, or Spark Structured Streaming
- Experience with data visualization tools such as Power BI, Tableau, or QuickSight
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.