Nomic- Senior Platform Engineer
HybridNew York City, New York, United States or New York, United States
Job Summary
Orchestrate agent rollouts across multi-account AWS and customer environments, managing deployment strategies, automated health checks, and pre-customer monitoring. Maintain inference performance and reliability at scale by overseeing GPU workloads, serving infrastructure, and cost-effective model execution. Build and evolve core infrastructure using Kubernetes, Terraform, and CI/CD pipelines while enforcing security posture through access controls, secrets management, and compliance. Design observability pipelines for traces, metrics, and logs to ensure SLOs are met as the platform scales to dozens of companies globally.
Required Qualifications
- 5+ years in infrastructure, DevOps, or SRE roles running cloud infrastructure in production
- Strong Kubernetes experience — deploying workloads, debugging real issues, working with operators and controllers
- Solid infrastructure-as-code skills — designing modules, managing state, reasoning about blast radius
- Strong software engineering fundamentals — you write and review production code in Python and/or TypeScript, not just infra configs
- Linux systems and networking fundamentals
- CI/CD pipeline design and maintenance
- A proactive orientation and genuine comfort owning a wide surface area
Desired Qualifications
- Terraform experience
- Observability platforms (Datadog, OpenTelemetry) — dashboards, trace/metric/log pipelines
- PostgreSQL operations — performance tuning, replica management
- ML/AI infrastructure — inference services, GPU workloads, model serving, eval pipelines
- Multi-tenant deployment patterns or per-customer isolation
- Experience building sandboxed execution environments or automated reliability systems
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.