KKR logo
KKRPosted 3 weeks ago

Data Engineer (SQL, PySpark, Snowflake, DBT, Redshift) - Associate

On-siteGurugram, Haryana, India

Full TimeEntry LevelLarge

Job Summary

Design, build, and maintain scalable, fault-tolerant data pipelines and ETL/ELT processes for structured, semi-structured, and unstructured data on AWS. Own end-to-end pipeline orchestration, including scheduling, dependency management, retries, and observability, while developing large-scale processing jobs using Apache Spark and AWS Glue. Model, load, and optimize data in Snowflake, manage Apache Iceberg lakehouse assets, and ensure data quality, integrity, and security across all systems. Integrate data from diverse sources including APIs, relational databases, and streaming feeds, then monitor and tune pipelines for performance and cost efficiency. Partner with data scientists and analysts to deliver reliable data products that power analytics across Insurance Systems. Mentor junior engineers and contribute to platform architecture and engineering best practices.

Required Qualifications

  • 5 to 8 years of hands-on data engineering experience building production data pipelines
  • Strong proficiency in Python for data engineering and automation
  • Advanced SQL and strong experience with relational databases (e.g., PostgreSQL, MySQL)
  • A data pipeline orchestrator is mandatory - proven experience operating a workflow orchestration tool (e.g., Apache Airflow, Dagster, or equivalent) in production
  • Hands-on experience with Apache Spark for large-scale distributed data processing
  • Production experience with Apache Iceberg (or an equivalent open table format) for lakehouse storage
  • Hands-on experience with AWS Glue (ETL jobs and the Glue Data Catalog)
  • Experience with Snowflake as a cloud data warehouse
  • Experience with data cataloging and governance using AWS Glue Data Catalog and Snowflake Horizon
  • Strong experience with AWS as the primary cloud platform, including S3, EMR, Lambda, Athena, Kinesis, Redshift, and IAM
  • Solid understanding of data modeling, data warehousing, and lakehouse architecture patterns
  • Working knowledge of REST APIs and data integration techniques
  • Strong problem-solving, analytical, and debugging skills
  • Strong communication and cross-functional collaboration skills
  • Ability to work in a fast-paced, agile environment
  • Self-driven with a proactive, ownership-oriented mindset
  • Ability to mentor peers and communicate technical concepts to non-technical stakeholders
  • Must be able to lift 50 lbs

Desired Qualifications

  • Experience with Dagster as a data pipeline orchestrator (strongly preferred)
  • Experience with containerization and orchestration (Docker, Kubernetes / Amazon EKS)
  • Exposure to CI/CD pipelines and infrastructure-as-code (e.g., Terraform, AWS CloudFormation/CDK)
  • Experience with streaming / real-time data (Kafka, Amazon Kinesis, Spark Structured Streaming)
  • Familiarity with data observability and quality frameworks (e.g., Great Expectations, dbt tests)
  • Knowledge of data governance, security, and compliance best practices
  • Experience within financial services or insurance data domains

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce