AI Engineering Lead
RemoteUnited States or Argentina
Job Summary
Lead end-to-end project delivery with clear governance and strong stakeholder communication, designing and building production-ready RAG systems, agentic frameworks, and LLM-powered solutions. Define AI system boundaries, communicate risks transparently to clients, and run structured experiments across prompts, retrievers, and models to ground decisions in evidence. Mentor junior engineers while designing evaluation frameworks, custom metrics, and go/no-go gates to ensure robust performance. Automate the full MLOps/LLMOps lifecycle, including tracking, versioning, deployment, and monitoring, while building scalable inference infrastructure and CI/CD pipelines. Identify and categorize model failure modes, design APIs and orchestration layers optimized for latency and cost, and lead feasibility assessments to select the right approach among prompting, fine-tuning, or classical ML.
Required Qualifications
- Lead end-to-end project delivery with clear governance and strong stakeholder communication
- Mentor junior engineers and contribute to proposals and new business initiatives
- Define what AI systems should and should not attempt, and communicate risks and tradeoffs transparently to clients
- Design and build RAG systems, agentic frameworks, and LLM-powered solutions robust enough for production
- Apply advanced prompt engineering techniques, including instruction design, few-shot sets, structured outputs, and tool/agent prompts
- Lead feasibility assessments to select the right approach among prompting, RAG, fine-tuning, or classical ML
- Design evaluation frameworks, including LLM-as-a-judge methods, custom metrics (recall@k, precision@k), and go/no-go gates
- Run structured experiments across prompts, retrievers, chunking strategies, and models, grounded in evidence rather than intuition
- Identify and categorize model failure modes, including hallucinations, retrieval misses, and instruction-following errors
- Build scalable inference infrastructure and CI/CD pipelines for AI/ML models
- Automate the full MLOps/LLMOps lifecycle, including tracking, versioning, deployment, monitoring, and retraining
- Design APIs, microservices, and orchestration layers optimized for latency, cost, and reliability
- Expert-level Python, strong Git practices, and experience with ML/LLM versioning
- Solid cloud experience across AWS, Azure, or GCP (Azure preferred), plus containerization and orchestration
- Hands-on RAG experience covering chunking, embeddings, retrieval, reranking, and evaluation
- Proven MLOps/LLMOps track record using tools such as MLflow, Weights & Biases, or similar
- Practical evaluation design skills, including metrics, dataset curation, and structured experimentation
- Experience with event-driven architectures, APIs, and microservices
- Strong communication skills, equally comfortable engaging engineering teams and senior stakeholders
- English: Advanced (required for effective communication with global teams)
- 6+ years of experience building and deploying AI solutions in production environments, with a strong track record across RAG, agentic systems, and MLOps/LLMOps
Desired Qualifications
- Preferred: experience with the Databricks MLOps platform, LLM fine-tuning, building agentic GenAI systems, Infrastructure as Code, security and observability for AI services, a classical ML background, and open-source contributions
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.