Site Reliability Engineer III - Machine Learning
On-siteWilmington, Delaware, United States
Job Summary
Guide the team in adopting site reliability engineering best practices by designing level designs and gaining peer consensus. Collaborate with software engineers to design, develop, test, and implement deployment approaches using automated continuous integration and delivery pipelines. Implement infrastructure, configuration, and network as code for applications and platforms while utilizing enterprise-authorized AI capabilities to accelerate incident triage, troubleshooting, and post-incident analysis. Resolve complex problems proactively using service level indicators and objectives to ensure the resilience and performance of critical banking applications. Apply AI to identify patterns indicating reliability risk or recurring toil, prioritizing reuse-first improvements tied to SLO outcomes.
Required Qualifications
- Formal training or certification on site reliability engineering concepts
- 3+ years applied experience
- Proficient in site reliability culture and principles
- Proficient in at least one programming language such as Python, Java/Spring Boot, and .Net
- Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data sensitivity
- Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements
- Proficient knowledge of software applications and technical processes within a given technical discipline (e.g., Cloud, AI, Android, etc.)
- Experience in observability such as white and black box monitoring, service level objective alerting, and telemetry collection
- Experience with continuous integration and continuous delivery tooling
- Familiarity with container and container orchestration
- Troubleshooting common networking technologies and issues
Desired Qualifications
- Experience orchestrating data/compute workflows (e.g., Airflow or AWS Step Functions)
- Familiarity with Databricks jobs, clusters, and workspace operations
- Experience using Python or REST APIs to interact with Databricks
- Infrastructure-as-Code experience
- Terraform Enterprise exposure
- GitOps and configuration management exposure (e.g., Argo CD/Flux patterns)
- Operational readiness automation tied to pull-request workflows
- Familiarity with OpenTelemetry or similar standards for metrics/logs/traces instrumentation and correlation
- AWS certification
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.