Senior Associate, Platform SRE Engineer, SRE&Governance, Group Technology
On-siteEast, Sichuan, People’s Republic of China
Job Summary
Develop monitoring and onboarding guidelines for applications using the observability platform stack, ensuring accurate data collection and implementing enterprise standards for tools like ELK, Grafana, and AppDynamics. Automate routine tasks and reporting processes via APIs and scripting to reduce manual effort, while identifying and resolving performance issues through detailed analysis of transaction traces, logs, and system metrics. Configure alerts to proactively detect problems, design custom dashboards for actionable insights, and integrate the observability stack with CI/CD pipelines. Collaborate with development and operations teams to manage user permissions, apply security controls, and perform application maintenance including patching and capacity planning. Ensure continuous uptime, adhere to bank standards for project tracking, and contribute to internal knowledge bases to promote team collaboration.
Required Qualifications
- University graduate (computer science or related field)
- Strong communication skills and ability to explain protocol and processes with team and management
- A passion for learning and using new technologies in the open-source communities
- Minimum 3 years of IT work experience
- Working knowledge in ELK Stack, Grafana, Open Telemetry (OTEL)
- Experience in triaging and troubleshooting application problems quickly in monitoring tools by using various techniques
- Knowledgeable and experienced in SRE (Site Reliability Engineering) practices covering monitoring, observability, performance management, automation, and resiliency
- Good understanding of Network routing, Load balancing and Networking protocols; a base knowledge of TCP/IP, with an understanding of HTTP and DNS
- Good problem diagnosis and creative problem-solving skills
Desired Qualifications
- Knowledge in Confluent Kafka, Prometheus & other APM tools (Dynatrace, Datadog, New Relic, Splunk)
- Knowledge in AI/ML capabilities to automate RCA's and shorter MTTR when issues arise
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.