Data Engineer – Data Platforms & AI Tooling (Remote, LATAM)
$64,800–$94,800 year
Remote
Job Summary
Build and maintain scalable batch and streaming data pipelines that ingest, transform, and prepare data for analytics, machine learning, and agentic workflows. Develop trusted, reusable data products and design models that enable reporting and secure AI consumption. Implement AI-facing integration layers, including MCP servers, and ensure data governance, privacy, and observability across all systems. Collaborate with cross-functional teams to support production-grade Agentic AI solutions and contribute to architecture decisions for platform scaling. This role is limited to candidates based in LATAM.
Required Qualifications
- 4+ years of experience in Data Engineering, Backend Engineering, or a related field
- Strong Python and SQL skills
- Experience building and operating production data pipelines
- Experience with cloud data platforms such as Snowflake, Databricks, BigQuery, Redshift, Microsoft Fabric, or similar
- Experience with dbt or comparable data transformation tools
- Experience with orchestration tools such as Airflow, Dagster, Prefect, Databricks Workflows, Azure Data Factory, or similar
- Experience with data modeling techniques such as dimensional modeling, SCDs, Data Vault, or comparable approaches
- Experience working with at least one major cloud provider (Azure, AWS, or GCP)
- Experience supporting, integrating, or developing AI-powered and Agentic AI solutions in production environments
- Familiarity with Retrieval-Augmented Generation (RAG) concepts, including embeddings, indexing, and retrieval workflows
- Experience with vector databases such as Pinecone, Weaviate, Qdrant, pgvector, or similar technologies
- Understanding of how AI agents consume data and interact with enterprise systems
- Experience building MCP servers or comparable agent-tool integration patterns
- Experience with CI/CD practices and version control workflows
- Familiarity with Docker and containerized environments
- Understanding of Infrastructure as Code (Terraform or similar tools)
- Experience implementing monitoring, telemetry, dashboards, and alerting for production systems
- Familiarity with OpenTelemetry (OTel) and observability best practices
- Active use of AI-assisted development tools such as GitHub Copilot, Cursor, Claude Code, or similar
- Excellent communication and collaboration skills
- Ability to work effectively with cross-functional engineering, platform, data, and AI teams
- Comfortable working in a fast-paced Agile (Kanban) environment
- Professional English proficiency
- Candidates based in LATAM
Desired Qualifications
- Experience with data quality, testing, monitoring, observability, and lineage across data pipelines and AI integrations
- Experience ensuring data governance, privacy, security, and access-control standards are applied consistently
- Experience preparing structured and unstructured data for analytics, machine learning, generative AI, and agentic workflows
- Experience contributing to architecture decisions and helping scale the organization's data platform for future growth
- Experience with agent evaluation and observability tools such as Langfuse, LangSmith, Promptfoo, DeepEval, LiteLLM, or similar
- Experience implementing data products that support analytics, business applications, and AI solutions
- Experience designing and maintaining data models that enable reporting, analytics, and AI consumption
- Experience building AI-facing integration layers, including MCP servers and similar interfaces that allow AI agents to securely access data and platform capabilities
- Collaborating with AI, platform, and infrastructure teams to support production-grade Agentic AI solutions
- Experience building and maintaining scalable batch and streaming data pipelines that ingest and transform data from applications, databases, APIs, event streams, and third-party systems
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.