Platform Operations Engineer, AVP
On-siteSingapore, Singapore
Job Summary
Administer, monitor, and troubleshoot enterprise cybersecurity platforms including SIEM, AI-SOC, telemetry routing, and Security Lakehouse solutions to ensure service stability and availability. Manage incident response, root cause analysis, and change execution while supporting detection engineering, threat hunting, and SOC operations. Maintain operational standards, governance, and ITIL-aligned procedures for compliance and continuous improvement across Splunk, Elastic, Cribl, Databricks, and other analytics technologies. Provide 24x7x365 operational support through an on-call rotation, collaborating with globally distributed teams to maintain platform reliability and service continuity.
Required Qualifications
- Strong experience administering and supporting enterprise-scale cyber security platforms including technologies such as Splunk, QRadar, Elastic, and similar security operations solutions
- Deep understanding of security monitoring, threat detection, incident response, correlation searches, notable event management, data models, and risk-based alerting methodologies
- Hands-on expertise to support telemetry and data pipeline solutions such as Cribl Stream, including data routing, transformation, filtering, enrichment, and optimization of machine data at scale
- Experience integrating security and operational telemetry into Databricks or similar data lakehouse platforms to support advanced analytics, reporting, machine learning, and AI-driven use cases
- Strong knowledge of cloud-based infrastructure and platform operations, preferably on AWS & Azure, combined with solid Linux/Unix administration, troubleshooting, and performance tuning skills
- Proficiency in scripting, automation, and version control practices to improve operational efficiency, automate repetitive tasks, enhance monitoring capabilities, and increase platform reliability
- Working knowledge of ServiceNow, Jira, and Confluence, with experience supporting incident, problem, change, service request, Agile delivery, and technical documentation processes
- Solid understanding of ITIL principles, operational best practices, service management disciplines, governance frameworks, and compliance requirements in a large enterprise environment
- Exceptional communication and stakeholder management skills, with the ability to effectively communicate technical concepts and operational risks to both technical and non-technical audiences
- Strong organizational and collaboration skills, with the ability to coordinate across geographically distributed teams, manage competing priorities, and successfully drive operational outcomes
- Experience supporting highly available, mission-critical platforms within complex enterprise environments, demonstrating a strong commitment to operational excellence, resiliency, and continuous improvement
- Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field
- 5+ years of platform engineering and platform/infrastructure operations experience
- 3+ years of experience with public cloud (AWS, Azure, OCI) & DevOps and ability to manage cloud-native platforms
- 5+ years of experience in supporting SIEM platforms like Splunk, Datadog, Qradar, Elastic, Cribl, Databricks and AI-SOC solutions
- Relevant certifications, such as AWS, Splunk, ITIL, CISSP, or related technology/security certifications
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.