Lead Enterprise Software Engineer – Application & IT Operations
On-siteChicago, Illinois, United States or Austin, Texas, United States
Job Summary
Own production reliability for critical applications by defining SLOs, error budgets, and capacity baselines while leading major incident response and driving data-driven root cause analysis. Direct release and change operations to enforce readiness gates and validate post-deployment health, and architect operational observability through dashboards, alert strategies, and log pipelines. Establish and continuously improve operational standards, guardrails, and runbooks; automate repetitive tasks to reduce toil. Partner with engineering on resiliency patterns and performance tuning, and plan capacity management, scaling strategies, and DR/BCP readiness. Champion security-by-default in operations, including secrets hygiene and vulnerability remediation. Mentor AppOps engineers and drive service reviews to publish operational KPIs and lead continuous improvement roadmaps. Maintain application environments, configuration baselines, and platform dependencies while ensuring consistency and compliance. Implement monitoring, logging, and observability using enterprise tools to ensure system transparency and rapid diagnosis. Develop and enhance runbooks, automate workflows, and improve operational efficiency. Track key reliability metrics, enforce operational standards, and drive continuous optimization to meet service commitments. Support change reviews, evaluate operational risks, and ensure compliance with change processes. Partner with Engineering, CloudOps, Security, and Compliance teams to resolve issues and enhance application resilience. Perform other duties as assigned by management.
Required Qualifications
- 8-10 Years
- Bachelor's degree in computer science, Information Systems, or a related field
- Ability to ensure 24x7 application reliability and operational excellence
- Manage end-to-end application lifecycle including deployments, configurations, and environment health
- Collaborate with engineering, CloudOps, and Security teams to ensure smooth operations
- Own operational KPIs such as uptime, MTTR, change success rate, and SLA/SLO adherence
- Perform release coordination, deployment validation, and post-release monitoring
- Lead incident response, communication, and escalation handling
- Participate in change management and risk assessments for all application changes
- Maintain runbooks, SOPs, and operational documentation
- Drive continuous improvement for operational workflows and process maturity
- Support audit, compliance, and security requirements for applications
- Advanced expertise in operating applications on Azure and/or AWS, including networking, load balancers, DNS, certificates, storage, and messaging services
- Strong knowledge of application operations in cloud environments (Azure/AWS)
- Hands-on with observability stacks (Datadog, Grafana/Prometheus, ELK/OpenSearch, Open Telemetry) and alert engineering
- Experience with incident management, RCA, and operational troubleshooting
- Strong practical understanding of CI/CD concepts and collaboration with release teams; experience validating releases in lower/production environments
- Familiarity with infrastructure components: load balancers, storage networking, DNS and certificates
- Proficiency in automation and scripting (PowerShell, Bash, Python) to build runbooks, health checks, and remediation workflows
- Experience with deployment strategies (blue/green, rolling, canary) and traffic management
- Security and compliance in operations: vulnerability remediation, secrets and key management, audit readiness
- Ability to interpret logs, metrics, traces, and performance data
- Experience managing multi-environment application lifecycles (Dev, QA, UAT, Prod)
- Infrastructure as Code (IaC): Terraform (modules, workspaces), Azure ARM/Bicep or AWS CloudFormation; policy-as-code and environment drift detection
- Vendor certifications preferred: Azure Administrator/Architect or AWS SysOps/DevOps Professional; ITIL Foundation (or higher)
- Terraform Associate/Professional (or equivalent IaC certification) preferred; SRE Foundation a plus
- Proven experience leading incident response, conducting RCAs, and implementing preventative controls
- Excellent communication, stakeholder management, and mentoring skills in global, fast-paced environments
- Strong understanding of Software Engineering Principals
- Industry recognized Kubernetes Certification
- Strong ownership mindset with a bias for automation, measurement, and continuous improvement
- Ability to translate technical risks and trade-offs into business language for decision-makers
- Must be able to lift 50 lbs
Desired Qualifications
- Vendor certifications preferred: Azure Administrator/Architect or AWS SysOps/DevOps Professional; ITIL Foundation (or higher)
- Terraform Associate/Professional (or equivalent IaC certification) preferred; SRE Foundation a plus
- Strong understanding of Software Engineering Principals
- Industry recognized Kubernetes Certification
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.