Site Reliability Engineer
On-siteSydney, New South Wales, Australia
Job Summary
Provide first and second level technical support for trading, reference data, surveillance, and clearing systems within the Markets Line of Business. Design and implement CI/CD pipelines, Infrastructure as Code, and automation for operational tasks and service recovery. Strengthen resilience through failover validation, disaster recovery exercises, and post-incident reviews to identify systemic causes and reduce recurrence. Manage observability practices across logs, metrics, and traces, ensuring alerts align with SLOs and operational risk. Support capacity planning, performance engineering, and production readiness assessments for market-critical services. Perform 24x7 on-call support, weekend installations, and upgrades while handling incident, problem, and change management.
Required Qualifications
- 5+ years of experience in a similar SRE role
- previous markets experience
- Strong understanding of incident response, post-incident review and problem management
- Experience with production readiness, release readiness and operational acceptance
- Experience with capacity, performance and resilience testing
- Strong knowledge of observability design, not only monitoring tools
- Experience supporting high-availability, distributed, business-critical platforms
- Strong scripting and automation skills using Python, PowerShell, shell scripting and scheduled automation tooling
- Demonstrated experience supporting critical applications using AWS Cloud Architecture covering EC2, S3, Lambda, RDS
- strong understanding of microservices, containerisation (Docker), Kubernetes administration, and CI/CD pipelines
- Experience setting up CloudWatch alerts, supporting observability tools such as Grafana, Prometheus, OpenTelemetry
- Strong database operational experience across Oracle and/or Microsoft SQL Server, including SQL scripting, performance tuning, indexing, query plans and production troubleshooting
- Strong Linux/Unix administration and troubleshooting experience across cloud, container and distributed application environments
- Excellent knowledge and technical skills on MS Windows Servers (2019-2022)
- Excellent troubleshooting, problem solving and root cause analysis skills
- Ability to communicate effectively with both technical and non-technical stakeholders
- Willingness to learn new technologies
- self-motivation with the ability to multi-task and prioritise work with minimal supervision
- AWS certification at Associate level or above
- Technical experience in distributed transactions, high availability, performance-critical systems
- Exposure to the standards, practices and procedures of the financial services industry
- Securities industry experience such as understanding of exchange-traded futures and options, derivatives markets, and the business processes related to clearing of member trades
- Strong networking troubleshooting across DNS, TLS, load balancers, firewalls and TCP/IP
- 24 x 7 on-call support (rostered)
- weekend/after-hour installations and upgrades
- incident, problem, release and change management
- BAU work within the team
- legally authorised to work in Australia on a permanent basis without any restrictions
Desired Qualifications
- reducing operational toil through automation
- improving observability across logs, metrics and traces
- supporting production readiness and non-functional testing
- driving resilience, capacity, incident learning and continuous reliability improvement across market-critical services
- reliability-first mindset to market-critical systems where availability, integrity, recoverability and controlled change are essential to maintaining confidence in ASX services
- Design and Implement: CI/CD pipeline, Infrastructure as Code (IaC) development and maintenance and automate operational tasks, deployments, and service recovery processes
- Contribute to production readiness assessments by reviewing observability, supportability, resilience, capacity, recovery procedures and operational documentation before release
- Support capacity planning and performance engineering by analysing utilisation trends, saturation signals, workload patterns and scalability risks across critical services
- Strengthen service resilience through failover validation, recovery testing, fault-tolerance reviews and participation in disaster recovery or operational resilience exercises
- Lead or contribute to post-incident reviews, identifying systemic causes, corrective actions and reliability improvements that reduce recurrence and improve recovery outcomes
- Design observability practices across logs, metrics and traces, ensuring alerts are actionable, dashboards reflect service health, and monitoring aligns to SLOs and operational risk
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.