HUD logo
HUDPosted 1 month ago

Platform Engineer

On-siteSan Francisco, California, United States

Full TimeEnterprise

Job Summary

Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services. Build and maintain AWS infrastructure with Terraform, Kubernetes, Helm, Docker, and networking while designing backend systems for scale, including capacity planning, autoscaling, and queueing. Define dashboards, alerts, logs, and on-call workflows to detect and resolve failures quickly. Construct reliable CI/CD pipelines, release automation, and environment management to improve developer productivity. Write clean code to automate systems and create internal tooling. Work within a ~15 person team focused on RL training data infrastructure for frontier AI agents, supported by Visa sponsorship and a 1-week work trial.

Required Qualifications

  • Have owned production cloud infrastructure for a high-availability, user-facing platform, with responsibility for uptime, performance, deployment safety, and cost
  • Have deep experience with AWS infrastructure and containerized systems
  • Have built or operated CI/CD, environment management, release automation, observability, alerting, and incident response systems
  • Have strong backend engineering judgment and can reason about service architecture, APIs, databases, async systems, queues, scaling limits, and production failure modes
  • Can write clean, maintainable code and apply strong software engineering judgment across product architecture, infrastructure, backend systems, and developer workflows

Desired Qualifications

  • Experience with tools like Terraform, Kubernetes/EKS, Docker, EC2, CodeBuild, ECR, S3, IAM, load balancers, networking, and secrets management
  • Experience operating infrastructure for data-heavy, ML/AI, workflow, marketplace, developer-tools, or enterprise platforms
  • Experience designing systems for bursty workloads, long-running jobs, sandboxed execution, distributed workers, or high-concurrency services
  • Experience reducing cloud spend through better architecture, autoscaling, workload placement, caching, cleanup systems, or observability
  • Experience building internal platforms or tools that make engineers faster without hiding too much complexity

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce