Senior Site Reliability Engineer
$160,000–$208,000 year
RemoteUnited States
Job Summary
Build systems for declarative application and infrastructure lifecycle management, including continuous deployment, continuous integration, Kubernetes cluster management, and service inventory. Prioritize and troubleshoot infrastructure issues to minimize downtime, while streamlining and automating delivery pipelines and database changes. Contribute to setting the direction of the Site Reliability Engineering team to ensure goals align with company objectives. Foster a collaborative, high-performance culture that promotes cross-disciplinary teamwork among technical leads, data scientists, and engineering professionals.
Required Qualifications
- 5+ years of programming experience
- proficiency in at least one of the following languages: Python, Go, or Shell Scripting
- in-depth knowledge of containerization technologies and orchestration, such as Docker, Containerd, and Kubernetes
- experience with CNCF-based technologies like Helm, gRPC, and Prometheus
- experience with public cloud platforms such as GCP, Azure, or AWS
- knowledgeable in networking fundamentals, including TCP/IP, UDP, firewalls, routing, DNS, and load balancing
- experience with Linux system administration
- solid understanding of Linux design principles
- understanding of key SRE concepts, such as monitoring, performance tuning, and automation
- ability to work autonomously with limited guidance
- ability to proactively identify and solve problems
- excellent communication and collaboration skills
- ability to work effectively with cross-functional teams
- ability to adapt to new challenges and evolving technologies
Desired Qualifications
- Kubernetes competency
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.