CoreWork Staffing logo
CoreWork StaffingPosted 1 month ago

Site Reliability Engineer

On-siteFlorida, United States or Dallas, Florida, United States

Full TimeSmall

Job Summary

Design, build, and maintain highly reliable distributed systems while defining Service Level Objectives and eliminating single points of failure. Manage cloud infrastructure across AWS, Azure, and GCP using Infrastructure as Code, and implement monitoring, logging, and alerting systems to analyze performance metrics and identify bottlenecks. Lead incident response and on-call rotations, performing root cause analysis and coordinating cross-functional teams during outages. Automate repetitive operational tasks and CI/CD pipelines to reduce manual overhead and improve deployment efficiency. Collaborate with software engineers and security teams to enforce best practices, ensure compliance, and advocate for reliability-focused architecture reviews.

Required Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field
  • 3+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or Systems Engineering
  • Strong experience with cloud platforms (AWS, Azure, and/or GCP)
  • Proficiency with Linux systems and networking fundamentals
  • Experience with containerization (Docker) and orchestration (Kubernetes)
  • Experience with CI/CD pipelines and automation tools
  • Strong scripting/programming skills (Python, Go, Bash, or similar)
  • Experience with monitoring and observability tools
  • Strong problem-solving and incident troubleshooting skills
  • Must currently reside in one of the approved locations listed above

Desired Qualifications

  • Experience with high-scale distributed systems
  • Knowledge of microservices and event-driven architectures
  • Familiarity with Infrastructure as Code (Terraform, Pulumi, CloudFormation)
  • Experience with SRE frameworks (Google SRE principles)
  • Knowledge of database systems (SQL/NoSQL) and performance tuning
  • Experience with service mesh technologies (Istio, Linkerd)
  • Familiarity with security practices in cloud-native environments
  • Experience in high-availability or mission-critical systems
  • Certifications in cloud or DevOps technologies

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce