AI/ML Engineer - Agentic
$136,500–$276,500 year
HybridSan Jose, California, United States
Job Summary
Design, build, and own a production-grade agentic orchestration platform, implementing scalable multi-agent workflows using frameworks like LangGraph. Architect and operate the MCP server infrastructure, managing inter-agent communication, tool registries, and domain isolation. Integrate enterprise-scale LLM services with streaming, structured outputs, and robust error handling, while building retrieval and memory services including RAG pipelines and hybrid search. Develop high-performance backend services using FastAPI and gRPC to power agent execution, and own observability, reliability, and cost visibility for non-deterministic systems. Manage cloud-native infrastructure on Kubernetes with automated CI/CD pipelines and resource optimization. This senior individual contributor role requires 4-7 years of experience and operates in a hybrid model with 2 days per week in an HPE office.
Required Qualifications
- Bachelor's degree in computer science, engineering, information systems, or closely related quantitative discipline
- 4-7 years' experience
- Production experience with agentic frameworks: LangGraph (preferred), Claude Agent SDK, or equivalent (not just prototypes)
- Deep understanding of multi-agent architectures: supervisor/worker patterns, hierarchical agent graphs, ReAct loops, ReWoo
- Hands-on with inter-agent communication protocols: MCP (Model Context Protocol), A2A, tool registry / server registry
- LLM API integration at scale: structured outputs, streaming, function/tool calling, error handling
- RAG pipeline design and optimization: chunking strategies, re-ranking, hybrid search
- Vector store experience: OpenSearch or equivalent
- Applied ML intuition: fine-tuning concepts, prompt engineering, evaluations, Qlora, PEFT
- Backend development: FastAPI, gRPC, Kafka, Redis, message queues
- Async System design: Python, API Design
- GraphQL and/or REST at enterprise scale
- Observability and monitoring for non-deterministic systems: LangFuse, Prometheus, or equivalent
- Kubernetes: deploying, scaling, and managing workloads (Deployments, Services, ConfigMaps, Secrets)
- Container image management: building, tagging, versioning, and pushing images via Docker; familiarity with a container registry (ECR, GCR, Docker Hub)
- CI/CD pipelines for automated build and deploy (GitHub Actions, Jenkins, ArgoCD, or similar)
- Resource management: CPU/memory limits, autoscaling (HPA/VPA), health probes
- Hybrid work arrangement: work on average 2 days per week from an HPE office
Desired Qualifications
- Master's degree
- Multi-tenant architecture awareness: rate limiting, auth, tenant isolation
- Knowledge base and cost optimization experience: AWS Bedrock, OpenSearch Serverless
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.