Site Reliability Engineer
RemoteUnited States
Job Summary
Lead the technical rearchitecting of the existing monolithic system into a resilient, cloud-native microservices architecture using domain-driven design principles. Drive cross-functional collaboration to define and implement the new system design while evaluating emerging tools, frameworks, and cloud services to enhance infrastructure. Manage container orchestration and management using Kubernetes across AWS and Azure to ensure extreme scalability, reliability, and high availability. Implement robust, highly resilient components and develop comprehensive monitoring, logging, and alerting mechanisms for optimal system performance. Drive the adoption of DevOps principles throughout the software development lifecycle to ensure seamless integration and continuous deployment. Mentor junior team members and provide technical guidance while staying current with industry trends in systems and cloud computing.
Required Qualifications
- U.S. Citizenship
- 100% remote within the US
- 7+ years of experience with cloud computing platforms
- Strong multi-cloud expertise with AWS and Azure
- 7+ years of experience in rearchitecting large-scale monolithic applications to cloud-native architectures
- Strong expertise in Kubernetes (K8S)
- Hands-on experience with AKS (Azure Kubernetes Service)
- Hands-on experience with EKS (Elastic Kubernetes Service)
- Strong experience with Cloud Networking
- Ability to design and resolve complex cloud networking architecture problems
- Expert knowledge of Terraform for infrastructure-as-code deployment and management
- Strong knowledge of security best practices for containers and Kubernetes clusters
- Bachelor's or Master's degree in Computer Science, Software Engineering, or a related field
Desired Qualifications
- Knowledge of load balancing algorithms
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.