Lambda logo
LambdaPosted 1 month ago

Senior Platform Engineer - Core Infrastructure

$230,000–$340,000 year

RemoteUnited States

Full TimeSenior LevelAssociates DegreeSmall

Job Summary

Architect, deploy, and operate Kubernetes clusters across AWS and Lambda's bare-metal datacenters. Build automation for cluster lifecycle management, including provisioning, upgrades, and scaling. Own reliability, performance, and security of production workloads while implementing observability, logging, and alerting. Partner with product teams to design scalable, cloud-native services and CI/CD pipelines, setting standards for resource management, networking, and RBAC. Lead incident response, root-cause analysis, and post-mortems for platform issues while mentoring engineers. Requires 5+ years in Platform, Infrastructure, or SRE roles with production Kubernetes experience. Full-time presence in San Francisco, San Jose, or Seattle offices four days per week.

Required Qualifications

  • 5+ years in Platform, Infrastructure, or SRE roles, including running Kubernetes in production at scale
  • Deep knowledge of Kubernetes internals and day-2 operations (upgrades, scaling, troubleshooting)
  • Strong with Helm, Kustomize, or similar, and GitOps-based delivery
  • Proficient with infrastructure-as-code (Terraform, Pulumi, or equivalent)
  • Solid grounding in networking, service meshes, and container runtimes
  • Hands-on with observability stacks (Prometheus, Grafana, OpenTelemetry)
  • Strong coding skills in Go or Python for automation and tooling
  • Practical security experience: network policies, secrets management, and image scanning
  • Presence in our San Francisco, San Jose, or Seattle office location 4 days per week
  • Lambda's designated work from home day is currently Tuesday

Desired Qualifications

  • Experience with multi-cluster, multi-cloud, or hybrid environments
  • Knowledge of GPU scheduling, HPC workloads, or ML/AI infrastructure
  • Experience with workflow orchestration / durable execution frameworks (Temporal, Cadence, or Argo Workflows)
  • Exposure to cost optimization and capacity planning for large clusters
  • Contributions to CNCF or Kubernetes open-source projects
  • CKA/CKS certification

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce