GRAIL logo
GRAILPosted 1 month ago

Staff Software Development Engineer - Enterprise AI Infrastructure - #4898

$169,000–$224,000 year

HybridDurham, North Carolina, United States or Menlo Park, California, United States

Full TimeSenior LevelLarge

Job Summary

Lead the end-to-end design, development, and monitoring of a scalable, governed enterprise AI platform leveraging Amazon EKS and AWS native services. Architect agentic AI workflows and multi-agent systems using advanced LLM orchestration techniques, while implementing secure integrations via the Model Context Protocol to connect internal systems and third-party SaaS applications. Enforce strict identity, authorization, and zero-trust token brokering flows using Okta, Auth0, and custom JWT authorizers, alongside deterministic policy controls to ensure role-based access and human-in-the-loop checks. Develop isolated, scalable containerized runtime environments for secure AI model execution and establish comprehensive audit trails for all AI interactions. Collaborate with Product, Security, Regulatory, and business stakeholders to translate enterprise requirements into compliant AI infrastructure solutions, while mentoring engineering teams and advancing the organization's enterprise AI strategy within a regulated medical device environment.

Required Qualifications

  • Bachelor's degree or equivalent in Computer Science, Software Engineering, Artificial Intelligence, Cloud Computing, or related field
  • Master's or PhD preferred
  • 8-12 years of relevant software development and cloud infrastructure experience with demonstrated technical leadership
  • Deep expertise in AWS cloud architecture and container orchestration, specifically with Amazon EKS, Kubernetes networking, network isolation (VPC, PrivateLink), IAM, KMS, and GenAI services (e.g., AWS Bedrock)
  • Proven experience in agentic AI development, building autonomous agents, and orchestrating LLM tool-calling workflows using frameworks like LangChain, LangGraph, AutoGen, or Claude Agent SDK
  • Hands-on experience implementing the Model Context Protocol (MCP) or building robust, governed API/tool integrations for LLMs
  • Strong background in identity and access management (IAM), OAuth, JWT, and integrating with enterprise IdPs (Okta, Auth0) for scoped, token-based authorization
  • Advanced proficiency in programming languages such as Python, TypeScript, or Go, and infrastructure-as-code tools (Terraform, AWS CDK)
  • Experience with vector databases, RAG (Retrieval-Augmented Generation) architectures, and row-level access controls (e.g., OpenSearch, FAISS, pgvector)
  • Proficiency with CI/CD pipelines, MLOps practices, Kubernetes ecosystem tools (e.g., Helm), containerization, and modern observability stacks
  • Demonstrated level of knowledge regarding applicable regulatory standards commensurate with the position's complexity and scope, contributing to organizational regulatory compliance
  • Cybersecurity principles, tools, and control frameworks (e.g., ISO 27001, NIST, SOC 2, HIPAA)
  • Operations within the regulated medical device environment (e.g., IVDD, IVDR, FDA 21 CFR 800 series, FDA 21 CFR Part 11)
  • AI governance, software validation, data integrity, and risk management principles applicable to regulated environments
  • Deep expertise in cloud infrastructure, containerized environments, agentic artificial intelligence, and secure distributed system design
  • Exceptional problem-solving and analytical skills with the ability to address ambiguous, high-impact technical challenges in AI orchestration and Kubernetes scaling
  • Strong leadership and influence skills, capable of driving alignment across engineering, security, regulatory, and business stakeholders
  • Excellent communication skills with the ability to explain complex LLM behaviors, infrastructure architectures, and security boundaries to technical and non-technical audiences
  • Proven mentoring and coaching capabilities that elevate cloud engineering and AI talent
  • Strong understanding of AI safety, prompt injection defenses, secure tool execution, and deterministic policy enforcement
  • Strategic thinking with the ability to balance long-term enterprise AI platform vision with near-term business delivery
  • High adaptability and intellectual curiosity regarding emerging agentic AI frameworks, MCP specifications, and cloud computing trends
  • Minimum of 60%, or 24 hours, of total work week on-site presence
  • Ability to work from GRAIL's office or from home
  • May require extended hours during major project deadlines, AI model deployments, production incidents, regulatory audits, or strategic initiatives

Desired Qualifications

  • Master's or PhD preferred

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce