Rakuten logo
RakutenPosted 1 month ago

Lead Site Reliability Engineer (Lead SRE) - Incentive Platform Department (INPD)

On-siteTokyo, Tokyo, Japan

Full TimeSenior LevelAssociates DegreeEnterprise

Job Summary

Lead Site Reliability Engineer (Lead SRE) responsible for defining technical direction for the stable operation and continuous improvement of mission-critical Rakuten incentive platform services. Lead the establishment of SLOs/SLAs, drive improvements in performance and latency, act as incident commander during outages, orchestrate RCA and systemic preventive measures, automate operational processes, and provide technical leadership and mentorship across SRE, product development, infrastructure, and security teams. Collaborates cross-functionally to align reliability goals, contribute to DevOps culture, and shape the 24/7 on-call rotation with refined runbooks.

Required Qualifications

  • Bachelor's degree in Computer Science or related field, or equivalent practical experience
  • More than 5 years of hands-on experience in SRE, infrastructure engineering, or a related field, with demonstrated technical leadership experience
  • Experience building and operating production systems in public cloud (AWS, GCP, Azure, etc.) or private cloud environments
  • Extensive experience designing, building, operating, and scaling Kubernetes environments
  • Deep knowledge and hands-on experience building and operating modern monitoring, alerting, and logging tools (e.g., Prometheus, Grafana, ELK Stack, Datadog)
  • In-depth knowledge of UNIX-like operating system internals and/or networking
  • Deep knowledge of IP network systems and protocols (TCP/IP, HTTP, etc.) and hands-on troubleshooting experience
  • Experience building automated workflows using CI/CD tools (e.g., Jenkins, CircleCI, GitLab, CI/CD)
  • Experience developing operational automation tools and scripts using scripting languages such as Shell, Python, etc.
  • Proven track record of leading production incident handling end-to-end (detection, triage, short-term / long-term fix, root cause analysis)
  • Experience in system performance tuning and capacity planning
  • Proficiency with Git and GitHub for version control and collaboration
  • Strong communication, negotiation, and collaboration skills to articulate complex technical issues and align with internal and external stakeholders

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce