Senior Data Platform Engineer (Data Lake and Catalog)
On-siteLondon, England, United Kingdom
Job Summary
Design, build, and evolve maintainable, scalable, and observable software for the Expedia Group data lake platform using JVM and Kotlin/Java. Lead technical design for services supporting the Lakehouse architecture, including Apache Iceberg and Hive integration, while owning end-to-end delivery and production operations. Mentor engineers through pairing and collaboration, break down complex problems into pragmatic solutions, and promote an SLO-driven culture focused on reliability and incident learnings. Partner with cross-functional stakeholders to clarify requirements and deliver scalable systems that improve data discoverability and trust. Contribute to open-source initiatives and stay current with technology trends to enhance the developer experience.
Required Qualifications
- Proven experience as a senior software engineer building and operating production systems at scale
- Strong knowledge of the JVM and server-side Kotlin/Java programming
- Solid understanding of the Hadoop ecosystem, including Spark, Hadoop, and Hive
- Strong experience with cloud platforms and data infrastructure, especially AWS, including EMR, S3, and Glue
- Experience working in Agile environments and contributing through code reviews, design discussions, and collaborative development
- Strong communication skills, with the ability to work effectively across technical and business stakeholders
- Ability to manage multiple priorities, make sound technical trade-offs, and deliver in a fast-paced environment
- Good understanding of architectural patterns and trade-offs, with confidence in recommending practical technical approaches
Desired Qualifications
- Strong passion for technology, engineering excellence, and solving complex engineering problems
- Experience delivering end-to-end solutions with a focus on scalability, reliability, and observability
- Experience contributing to open-source communities with tools like Circus Train, a Hive dataset replication tool, and Waggle Dance, a federation service that allows access to data lake tables across multiple catalogs
- Experience with Hive support to Apache Iceberg, a high-performance open table format for large analytic datasets that underpins many modern Lakehouse architectures
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.