SRE Software Engineer III
On-siteJersey City, New Jersey, United States
Job Summary
Design and implement automated continuous integration and continuous delivery pipelines to improve release quality, speed, and repeatability. Develop software solutions that strengthen availability, reliability, scalability, and operational readiness while partnering with stakeholders to troubleshoot complex, multi-system problems. Define service level indicators and objectives to proactively identify risk, and advance observability by improving telemetry, dashboards, and alerting to shorten time-to-detect and time-to-recover. Improve operational excellence through runbooks, automation, and continuous improvement actions that reduce toil and production risk, leveraging enterprise-authorized AI coding assist tools to enhance code quality and productivity.
Required Qualifications
- Formal training or certification on software engineering concepts
- 3+ years applied experience
- Proficiency in site reliability culture and principles
- Proficiency in at least one programming language such as Python, Java/Spring Boot, or .NET
- Experience building observability solutions, including telemetry collection and service level objective-based alerting using tools such as Grafana, Dynatrace, Prometheus, Datadog, or Splunk
- Experience with continuous integration and continuous delivery tools such as Jenkins, GitLab, or Terraform
- Familiarity with containers and orchestration technologies such as Docker, Kubernetes, or Amazon Elastic Container Service
- Working knowledge of diagnosing and troubleshooting common networking concepts and issues in distributed systems
- Ability to proactively remove blockers, learn new technologies, and apply new approaches to improve delivery outcomes
- Working knowledge of using enterprise-authorized AI-assisted software development tools within the work environment (e.g., for coding, test creation, troubleshooting, or documentation) with demonstrated ability to critically evaluate, validate, and refine AI-generated outputs for correctness, performance, and security
- Understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; ability to guide peers on safe and effective usage within team practices
Desired Qualifications
- Experience designing reliability improvements using error budgets, capacity planning, and resilience patterns (e.g., rate limiting, backpressure, graceful degradation)
- Experience improving release safety with progressive delivery practices (e.g., canary deployments, feature flags, automated rollbacks)
- Experience implementing automated quality gates (unit, integration, performance, and security testing) within delivery pipelines
- Exposure to incident management practices, post-incident reviews, and implementing measurable remediation actions to prevent recurrence
- Familiarity with infrastructure-as-code and standardized environment provisioning to improve consistency and auditability
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.