Blend logo
BlendPosted 3 weeks ago

AI Engineering Lead

RemoteUruguay

Full TimeSenior LevelMedium

Job Summary

Lead end-to-end project delivery with clear governance and strong stakeholder communication while mentoring junior engineers and contributing to proposals. Design and build production-ready RAG systems, agentic frameworks, and LLM-powered solutions by applying advanced prompt engineering techniques and running structured experiments grounded in evidence. Define AI system boundaries, communicate risks transparently to clients, and design evaluation frameworks including LLM-as-a-judge methods and custom metrics. Build scalable inference infrastructure, automate the full MLOps/LLMOps lifecycle, and design APIs optimized for latency and reliability. Identify model failure modes and lead feasibility assessments to select the right approach among prompting, RAG, fine-tuning, or classical ML. Requires 6+ years of experience, expert-level Python, and hands-on cloud experience across AWS, Azure, or GCP.

Required Qualifications

  • Lead end-to-end project delivery with clear governance and strong stakeholder communication
  • Mentor junior engineers and contribute to proposals and new business initiatives
  • Define what AI systems should and should not attempt, and communicate risks and tradeoffs transparently to clients
  • Design and build RAG systems, agentic frameworks, and LLM-powered solutions robust enough for production
  • Apply advanced prompt engineering techniques, including instruction design, few-shot sets, structured outputs, and tool/agent prompts
  • Lead feasibility assessments to select the right approach among prompting, RAG, fine-tuning, or classical ML
  • Design evaluation frameworks, including LLM-as-a-judge methods, custom metrics (recall@k, precision@k), and go/no-go gates
  • Run structured experiments across prompts, retrievers, chunking strategies, and models, grounded in evidence rather than intuition
  • Identify and categorize model failure modes, including hallucinations, retrieval misses, and instruction-following errors
  • Build scalable inference infrastructure and CI/CD pipelines for AI/ML models
  • Automate the full MLOps/LLMOps lifecycle, including tracking, versioning, deployment, monitoring, and retraining
  • Design APIs, microservices, and orchestration layers optimized for latency, cost, and reliability
  • Expert-level Python, strong Git practices, and experience with ML/LLM versioning
  • Solid cloud experience across AWS, Azure, or GCP (Azure preferred), plus containerization and orchestration
  • Hands-on RAG experience covering chunking, embeddings, retrieval, reranking, and evaluation
  • Proven MLOps/LLMOps track record using tools such as MLflow, Weights & Biases, or similar
  • Practical evaluation design skills, including metrics, dataset curation, and structured experimentation
  • Experience with event-driven architectures, APIs, and microservices
  • Strong communication skills, equally comfortable engaging engineering teams and senior stakeholders
  • English: Advanced (required for effective communication with global teams)
  • 6+ years of experience building and deploying AI solutions in production environments, with a strong track record across RAG, agentic systems, and MLOps/LLMOps

Desired Qualifications

  • Preferred: experience with the Databricks MLOps platform, LLM fine-tuning, building agentic GenAI systems, Infrastructure as Code, security and observability for AI services, a classical ML background, and open-source contributions

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce