AI Platform Engineer
$163,900–$215,000 year
HybridNew York City, New York, United States or Boston, Massachusetts, United States
Job Summary
Define architectural direction for the AI platform, leading design on the LLM gateway, multi-tenant compute isolation, and model serving infrastructure. Write ADRs to establish engineering standards and serve as the first point of technical escalation for hard decisions. Run design reviews, provide deep feedback on pull requests, and pair with engineers on complex problems. Own end-to-end execution of major platform initiatives, managing technical risk and driving projects to production. Lead reliability strategy by defining SLOs, shaping observability, and conducting incident reviews. Design governance and compliance controls for data residency, access management, and audit logging to satisfy enterprise requirements. Drive cross-team alignment with AI engineering, product, and cloud engineering teams to communicate technical trade-offs.
Required Qualifications
- 5+ years in platform, infrastructure, or SRE, with a track record as a technical lead or staff-level IC with team-wide technical scope
- Certified Kubernetes Administrator (CKA), Certified Kubernetes Application Developer (CKAD) or equivalent AWS Certifications
- 3+ years experience in cloud-native architecture: Kubernetes at scale, managed cloud services, networking, identity federation, and multi-tenancy patterns across AWS, GCP, or Azure
- 3+ years experience of proven ownership of complex, multi-month platform initiatives—driven from whiteboard to production, managing ambiguity and technical risk throughout
Desired Qualifications
- Strong IaC and GitOps fluency: Terraform or Pulumi, ArgoCD, with experience standardizing platform tooling and deployment patterns across engineering teams
- Hands-on experience with AI/ML infrastructure: model serving, inference pipelines, GPU resource management, or LLM integration patterns
- Experience mentoring engineers and influencing technical direction across teams
- Strong written communication: clear design docs, useful ADRs, and the ability to explain architectural decisions to both engineers and non-technical stakeholders
- LLM serving at scale—vLLM, Triton, Ray Serve—or experience with AI gateway design patterns
- Experience building an internal developer platform (IDP) from scratch, with a product mindset that obsesses over internal developer experience
- Strong grasp of AI safety, model evaluation, and governance frameworks
- Ability to lead through technical credibility—influencing design decisions, driving alignment, and raising quality without formal authority
- FinOps or GPU cost optimization experience across large inference workloads
- Open-source contributions to platform or ML infrastructure tooling
- Comfort with ambiguity: energized by undefined problem spaces, with a habit of building clarity and shared context where there isn't any
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.