Poolside logo
PoolsidePosted 1 month ago

Member of Engineering (Compute)

Remote

Full TimeSmall

Job Summary

Design and develop internal scheduling systems to maximize GPU utilization across the company. Build APIs and tooling to manage GPU workload lifecycles and troubleshoot failures. Improve the inference control plane to accelerate model deployment and request serving. Collaborate with research, scalability, and infrastructure teams to stabilize large-scale fault-tolerant training and ensure full node utilization. Optimize GPU utilization and deliver a stable, scalable inference serving stack for researchers. This role supports the compute team's mission to improve research velocity and enable frontier model development through intelligence systems built for software development.

Required Qualifications

  • Strong programming skills in Go, or other similar languages
  • Strong systems engineering background: distributed systems, schedulers, control planes, or high-throughput data planes
  • Production experience with Kubernetes internals — controllers, informers, operators — not just deploying to it
  • Bias toward observability and debuggability: building a system that is easy to navigate when debugging production issues

Desired Qualifications

  • Plus: experience in systems serving large scale inference requests

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce