Technical Site Reliability Engineer
On-siteAbu Dhabi, Abu Dhabi, United Arab Emirates
Job Summary
Maintain the simulation software stack, including installation, configuration, updates, and version management across the Simulation Center's tools. Own the underlying infrastructure for compute, networking, storage, and environment configuration, ensuring systems remain provisioned, patched, and performant. Build and maintain a post-release test suite to automate regression and smoke tests, catching integration issues immediately after software releases. Forecast, diagnose, and eliminate failure modes by root-causing errors and implementing guardrails to prevent recurrence. Partner with development teams to review changes for reliability risks and surface concerns early. Monitor overall system health by instrumenting the environment, triaging issues, and escalating clearly when problems exceed your scope. Document operational knowledge through runbooks, known issues, and release validation results to ensure the Simulation Center's capabilities are not held in one person's head.
Required Qualifications
- Proficiency in Python for automation, tooling, and test development.
- Working knowledge of C++; enough to read, debug, build, and trace issues in the simulation codebase.
- Solid general networking fundamentals: TCP/IP, UDP, multicast, DNS, routing, firewalls, and the ability to diagnose latency, packet loss, and connectivity problems across distributed systems.
- Experience with project management, issue tracking, bug triage, and coordinating work across engineering teams.
- Demonstrated experience maintaining production or production-adjacent systems, including troubleshooting under time pressure.
- Strong written and verbal communication; you can escalate an issue, explain a root cause, and write a runbook someone else can follow.
- Eligibility to pass the security and background check requirements for sensitive information systems.
- This position is in Abu Dhabi, UAE, with initial position hiring occurring in London, UK.
- Candidate must be willing to relocate to the facility upon completion.
Desired Qualifications
- Experience with modeling and simulation, wargaming, or distributed simulation standards (DIS, HLA, TENA) and platforms such as AFSIM, VBS, or similar.
- Test automation and CI/CD experience; building automated validation pipelines, not just running them.
- Infrastructure-as-code and configuration management (Terraform, Ansible, Docker, Kubernetes).
- On-prem and cloud deployment experience.
- Observability tooling: Prometheus, Grafana, ELK, or equivalent.
- Linux systems administration depth; comfort in mixed Linux/Windows environments.
- Prior work in a defense, aerospace, or classified environment.
- Active security clearance.
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.