Senior Data Engineer
On-siteHong Kong, Hong Kong
Job Summary
Build, optimize, and manage the data lake and associated processing frameworks to ensure reliable, scalable data ingestion from diverse sources. Architect and implement robust ETL pipelines using orchestration tools to move data through various stages into the lake. Structure data schemas and design models to support analytics, reporting, and machine learning initiatives while integrating new data streams with product teams. Evaluate emerging tools and frameworks to continuously improve platform performance and capabilities. This senior role focuses on owning end-to-end projects within a collaborative, high-performing team environment.
Required Qualifications
- Solid understanding of data lake architectures
- Experience working with columnar big data databases (e.g., Athena, Redshift, Vertica, Hive/Hadoop)
- Familiarity with Iceberg
- Proven ability to design and implement ETL pipelines and manage workflows using tools such as Apache Airflow, Luigi, or AWS Batch
- Strong Python skills
- Hands-on experience using relevant libraries (e.g., boto3, pandas, pytest)
- PySpark experience
- Practical experience with cloud environments—particularly AWS (Glue, EMR, EC2, S3, Lambda, IAM, CloudWatch) or equivalent platforms
- Familiarity with Docker and Kubernetes (or AWS ECS/EKS) for deployment and scaling
- Experience with CI/CD tools (e.g., CircleCI, Jenkins, AWS CodePipeline) and solid git practices (branching strategies, collaboration workflows)
- Comfortable working within Agile/Lean frameworks such as Scrum or Kanban
- Communicate effectively in English
- Adaptable
- Capable of owning projects from end-to-end
- Thrive in collaborative, high-performing team environments
- Eager to learn, grow, and fill any gaps on the job
- Approximately 3+ years of relevant data engineering experience
Desired Qualifications
- Familiarity with Iceberg is a significant advantage
- PySpark experience is highly desirable
- Distributed messaging and streaming systems (Kafka, Pulsar, RabbitMQ)
- Streaming processing frameworks (Spark Streaming, Apache Beam, Apache Flink)
- Metadata catalogue and lineage tools (Amundsen, Apache Atlas, Alation)
- JVM languages (Java, Scala, Kotlin) and related frameworks
- RDBMS/NoSQL databases (PostgreSQL, MySQL, DynamoDB, Redis)
- BI tools (Tableau, Looker, PowerBI, QuickSight)
- Logging and monitoring stacks (ELK, Datadog, Prometheus, Grafana)
- Data privacy and security concepts (encryption, tokenization, Apache Ranger)
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.