Swift logo
SwiftPosted 1 month ago

Lead Site Reliability Engineer

On-siteKuala Lumpur, Kuala Lumpur, Malaysia

Full TimeSenior LevelLarge

Job Summary

Lead end-to-end delivery pipelines, capacity planning, and architecture design while developing automation scripts and infrastructure as code to improve system reliability. Analyze production issues, identify root causes, and implement long-term improvements through monitoring, observability setups, and architectural enhancements. Automate multi-tenant infrastructure deployment, maintain dashboards, and conduct blameless postmortems during on-call rotations. Collaborate with developers and stakeholders to troubleshoot systems, provide guidance to junior members, and organize efficient handovers through high-quality documentation.

Required Qualifications

  • Bachelor's/master's degree in engineering, Computer Science, IT, or equivalent experience
  • Minimum 10 years of experience in Site Reliability Engineering or software development within an international company
  • Minimum 2 years of experience leading project
  • Familiarity or experience with data ingestion with big data technologies (Elastic Search, Logstash, Kibana and kafka)
  • Experience with CICD development & deployment tools such as Maven, Jenkins, Nexus, Git, and Docker
  • Proficiency in Linux OS
  • Proficiency in scripting and automation (e.g. Python, PowerShell, YAML) with the ability to develop tools and infrastructure as code (Preferably Ansible, Terraform, Kubernetes, OpenShift)
  • Understanding of distributed systems and microservices architectures, including REST and SOAP APIs
  • Experience working within an Agile-driven environment
  • Practical experience in building metrics for data-driven reporting
  • Strong interpersonal skills with a customer-centric mindset and ability to work effectively across diverse cultures
  • Proven ability to collaborate with both local and remote teams across different time zones

Desired Qualifications

  • Hands-on experience with ITIL processes, including Incident, Problem, and Continual Improvement, is an advantage
  • Familiarity with or experience in managing VM hosts using vCenter is an advantage

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce