Nexthink logo
NexthinkPosted 2 months ago

Senior Site Reliability Engineer

HybridMadrid, Madrid, Spain

Full TimeSenior LevelMedium

Job Summary

Implement and manage cloud-native systems on AWS, operating Kubernetes clusters and deployment pipelines to support rapid delivery cycles. Design, build, and maintain infrastructure for a multi-tenant SaaS platform using Terraform, while defining SLOs, SLAs, and error budgets to proactively address availability and performance issues. Monitor applications and participate in a shared on-call rotation, acting as Incident Commander to coordinate cross-team responses and drive improvements to Mean Time to Detect and Recovery. This role supports over 50 Product Engineering teams and the Technical Platform, Security, and Architecture teams to enhance system reliability and scalability. You will work closely with software engineers to embed observability and fault tolerance principles into service design, automating runbooks and health checks to ensure safe, fast releases. Join Nexthink's SRE organization to strengthen the infrastructure powering digital employee experience management for 18 million employees.

Required Qualifications

  • Minimum Bachelor's degree in Computer Science or equivalent practical experience
  • 5+ years of experience as a Site Reliability Engineer or Platform Engineer with strong knowledge of software development best practices
  • Strong hands-on experience with public cloud services (AWS, GCP, Azure) and supporting SaaS product
  • Strong programming or scripting skills (e.g., Python, Go, Bash...), and experience with infrastructure-as-code (e.g. Terraform)
  • Proficiency with Kubernetes, container-based deployment (e.g., Docker) and related ecosystems (e.g., Helm)
  • Experience supporting multi-tenant microservices architectures
  • Experience with CI/CD pipelines & tools (e.g., Jenkins, GitHub Actions, GitLab CI, FluxCD, Crossplane)
  • Experience with managing monitoring solutions (e.g. Datadog)
  • Comfortable participating in a rotating on-call schedule, managing critical incidents, and leading post-incident reviews
  • At ease with operating and managing production systems, striking the right balance between urgency and methodology
  • Strong system-level troubleshooting skills and a proactive mindset toward incident prevention
  • Deep understanding of Linux systems, networking, and common troubleshooting practices
  • Solid understanding of the network stack (e.g., TCP/IP, VPN, etc.), cloud architectures (VPC, subnets, firewalls, load balancers), service mesh (e.g., Istio) and storage (e.g., S3, EBS, etc)
  • Knowledge of zero-downtime deployment strategies, blue/green and canary releases
  • Exposure to compliance standards such as SOC 2, ISO 27001, or HIPAA
  • Experience with chaos engineering or resilience testing practices
  • Excellent problem-solving skills, collaborative mindset, and a strong grasp of agile, iterative development
  • Self-driven, highly organised, and capable of independently managing priorities
  • Curiosity to learn new things and discover new technologies
  • Strong communication, presentation, and team collaboration skills
  • Excellent written and verbal skills in English

Desired Qualifications

  • FedRAMP experience is a big plus
  • The prior experience with any of the above-mentioned tools is a bonus, but not a must!

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce