AI Engineering Lead
RemoteUruguay
Job Summary
Lead end-to-end project delivery with clear governance and strong stakeholder communication while mentoring junior engineers and contributing to proposals. Design and build production-ready RAG systems, agentic frameworks, and LLM-powered solutions by applying advanced prompt engineering techniques and running structured experiments grounded in evidence. Define AI system boundaries, communicate risks transparently to clients, and design evaluation frameworks including LLM-as-a-judge methods and custom metrics. Build scalable inference infrastructure, automate the full MLOps/LLMOps lifecycle, and design APIs optimized for latency and reliability. Identify model failure modes and lead feasibility assessments to select the right approach among prompting, RAG, fine-tuning, or classical ML. Requires 6+ years of experience, expert-level Python, and hands-on cloud experience across AWS, Azure, or GCP.
Required Qualifications
- Lead end-to-end project delivery with clear governance and strong stakeholder communication
- Mentor junior engineers and contribute to proposals and new business initiatives
- Define what AI systems should and should not attempt, and communicate risks and tradeoffs transparently to clients
- Design and build RAG systems, agentic frameworks, and LLM-powered solutions robust enough for production
- Apply advanced prompt engineering techniques, including instruction design, few-shot sets, structured outputs, and tool/agent prompts
- Lead feasibility assessments to select the right approach among prompting, RAG, fine-tuning, or classical ML
- Design evaluation frameworks, including LLM-as-a-judge methods, custom metrics (recall@k, precision@k), and go/no-go gates
- Run structured experiments across prompts, retrievers, chunking strategies, and models, grounded in evidence rather than intuition
- Identify and categorize model failure modes, including hallucinations, retrieval misses, and instruction-following errors
- Build scalable inference infrastructure and CI/CD pipelines for AI/ML models
- Automate the full MLOps/LLMOps lifecycle, including tracking, versioning, deployment, monitoring, and retraining
- Design APIs, microservices, and orchestration layers optimized for latency, cost, and reliability
- Expert-level Python, strong Git practices, and experience with ML/LLM versioning
- Solid cloud experience across AWS, Azure, or GCP (Azure preferred), plus containerization and orchestration
- Hands-on RAG experience covering chunking, embeddings, retrieval, reranking, and evaluation
- Proven MLOps/LLMOps track record using tools such as MLflow, Weights & Biases, or similar
- Practical evaluation design skills, including metrics, dataset curation, and structured experimentation
- Experience with event-driven architectures, APIs, and microservices
- Strong communication skills, equally comfortable engaging engineering teams and senior stakeholders
- English: Advanced (required for effective communication with global teams)
- 6+ years of experience building and deploying AI solutions in production environments, with a strong track record across RAG, agentic systems, and MLOps/LLMOps
Desired Qualifications
- Preferred: experience with the Databricks MLOps platform, LLM fine-tuning, building agentic GenAI systems, Infrastructure as Code, security and observability for AI services, a classical ML background, and open-source contributions
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.