Weights & Biases logo
Weights & BiasesPosted 1 month ago

Senior Software Engineer, Production Engineering (Cloud & On-Prem) - W&B

$139,000–$185,000 year

On-siteLivingston, New Jersey, United States

Full TimeSenior LevelMedium

Job Summary

Design, build, and operate critical reliability and infrastructure services across AWS, GCP, Azure, and on-premises environments. Improve error attribution and alert routing to automatically route pages to owning teams. Own observability patterns, SLI/SLO frameworks, and dashboards providing visibility from code change through production. Build release-safety systems including canary deployments, smoke tests, and staged rollouts. Advance the incident and on-call program by managing tooling, rotations, and operational readiness reviews. Reduce on-call burden through architecture and automation while participating in rotations yourself. Provision infrastructure with Terraform and drive manual operations toward automated, repeatable code. Break down large, ambiguous operational problems into shippable engineering work. Provide technical leadership and mentorship alongside senior and principal engineers to influence platform direction.

Required Qualifications

  • Extensive engineering experience designing, building, deploying, and operating critical production infrastructure and services across a large enterprise
  • Expert in one or more major public clouds (AWS, GCP, or Azure), with real experience operating on-premises or hybrid environments, and in multi-cloud environments with multi-account strategies
  • Strong skills in infrastructure-as-code, automation, and configuration management (Terraform, CloudFormation, CDK, Ansible, or similar)
  • Hands-on expertise with Kubernetes and containerized workloads in production
  • Proficient in a systems/scripting language (Go, Python, Bash, or similar) and comfortable writing tooling and automation, not just operating it
  • Deep experience with CI/CD (GitHub Actions) and observability / monitoring tooling (Prometheus, Grafana, Datadog, etc.)
  • Proven ability to evolve designs to meet increasingly challenging scale, reliability, and performance requirements
  • Comfortable owning on-call for services you build, and passionate about reducing on-call burden through architecture and automation rather than heroics
  • Proven ability to provide technical leadership, work effectively with senior and principal engineers, influence technical direction, and contribute to the success of cross-functional stakeholders
  • Must be a U.S. person (U.S. citizen, national, lawful permanent resident, refugee, or asylee) or eligible to access export controlled information without a required export authorization, or eligible and reasonably likely to obtain the required export authorization

Desired Qualifications

  • Experience building and operating SaaS products
  • Experience building and operating reliability, incident, or developer-productivity platforms (service catalogs / Backstage, PagerDuty tooling, SLO frameworks)
  • Experience operating stateful systems and databases in production (PostgreSQL, MySQL, ClickHouse, etc.)
  • Cloud certifications (AWS Solutions Architect, Azure Architect, etc.)
  • Familiarity with security and access-management best practices in hybrid environments

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce