Principal Platform Engineer
$190,000–$225,000 year
RemoteUnited States
Job Summary
Design, build, and operate the core cloud infrastructure on AWS, including compute, networking, identity, and container orchestration. Own the CI/CD platform, observability stack, and incident response tooling while driving infrastructure as code practices across the organization. Partner with security, finance, and engineering teams to manage secrets, cost visibility, and self-service developer tooling that enables autonomous service deployment. Lead major incident response, maintain reliability standards, and author documentation and runbooks to ensure platform adoption without handholding. Work normal business hours in the Eastern US time zone.
Required Qualifications
- BS (or higher, e.g., MS or Ph.D.) in Computer Science or a related technical field involving coding, or equivalent technical experience
- 8+ years of platform, infrastructure, or site reliability engineering experience, with demonstrated Staff, Principal, or Architect-level scope
- Deep, hands-on experience designing and operating CI/CD pipelines for high-velocity engineering organizations, including artifact management, environment promotion, and progressive rollout
- Strong AWS background, comfortable down to the IAM, networking, and container orchestration layers, including production Kubernetes experience
- Experience with GitHub and GitHub actions
- Proven track record building internal developer platforms or tools that other engineering teams adopted by choice, not by mandate
- Hands-on coding fluency in Python, Go, or TypeScript
- Comfortable operating in a polyglot environment
- Practical experience with infrastructure as code (such as Terraform, CloudFormation, or Pulumi) and modern observability tooling (such as Prometheus, New Relic, Grafana, Datadog, or Open Telemetry)
- Comfortable owning the cost and reliability conversation with both engineering leadership and finance partners
- Strong written communication and a bias toward documentation, runbooks, and clear interfaces
- Proven analytical thinking and problem-solving skills
- Excellent communication skills, both verbal and written
- Work normal day business hours in the Eastern US time zone
Desired Qualifications
- Experience supporting infrastructure for AI or LLM-powered workloads, including GPU provisioning, model serving, or agent runtime infrastructure
- Background in healthcare, PBM, pharmacy, or another regulated data environment
- Kubernetes operator experience or comfort building custom controllers
- FinOps experience, particularly attributing infrastructure spend to features, teams, or business units
- Familiarity with distributed systems tooling such as Kafka, gRPC, or Dragonfly and Redis
- Experience with Agile development methodologies, preferably both Scrum and Kanban
- Experience with code scanning, security, and quality checks as part of the CICD
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.