Lead Data Engineer ID71008
On-siteBogotá, Bogota D.C., Colombia
Job Summary
Design and own ETL pipelines that extract, transform, and validate data from internal databases and external APIs at scale. Make architectural decisions on partitioning, file formats, schema strategy, and near-real-time processing for large-volume, OLAP-oriented data systems built on an object-storage data lake. Own the design of scheduled batch workflows using DAG-based orchestration with Airflow, driving architectural conversations while enforcing code quality standards and guiding senior developers. This role requires high autonomy to define underspecified problems, gather context, and shape technical approaches before implementation. You will drive the AWS data stack, including S3-backed data lakes, Athena, and EKS, partnering directly with DevOps to clarify requirements.
Required Qualifications
- 7+ years of engineering experience
- proven track record designing and implementing ETL pipelines
- making architectural decisions for large-volume data systems
- Hands-on experience with OLAP-style analytical data architecture
- comparable experience with Athena, Trino/Presto, BigQuery, Snowflake, Spark SQL, ClickHouse, or similar
- OLAP depth
- hands-on experience designing against a data lake sitting on object storage
- S3 or equivalent
- serverless engine
- partitioning strategy
- file formats (Parquet/ORC)
- cost/performance tradeoffs
- Deep familiarity with DAG-style workflow definition and triggering
- substantial prior Airflow experience
- or enough depth in a comparable orchestrator (Dagster, Prefect, Luigi, Step Functions)
- practical comfort across the AWS data stack
- S3-backed data lake
- serverless query engines (Athena or equivalent)
- EKS/Kubernetes
- ability to drive infrastructure conversations with DevOps
- Backend proficiency in Python
- FastAPI or Flask
- Comfortable with REST and GraphQL
- Docker
- PostgreSQL
- Mac/Linux terminal-centric environments
- Practical, hands-on use of AI-assisted development tools
- critical judgment to challenge AI output when it compromises long-term maintainability
- leadership presence to set the standard for how the team uses AI tooling responsibly
- ability to hold and defend a technical opinion
- challenging a stakeholder's or a tool's proposed quick fix with sound reasoning
- pursuit of a solution that scales and is maintainable long-term
- pragmatic enough to ship
- Comfort with ambiguity
- work is frequently ad hoc and underspecified
- defining the problem
- gathering context
- identifying constraints
- framing the work
- Upper-Intermediate English level
Desired Qualifications
- Direct production experience with Athena specifically
- Working knowledge of TypeScript/React
- enough to guide integration and review frontend-adjacent PRs
- Production AI features using AWS Bedrock, LangChain, Pydantic AI, or similar
- Monorepo tooling (Nx)
- modern package managers (Poetry, UV, Yarn)
- Redis/caching layers
- SageMaker
- Experience with marketing data structures
- campaign management APIs
- digital advertising metrics
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.