Senior DevOps Engineer - Azure AKS & Platform Security
On-siteReading, Pennsylvania, United States
Job Summary
Design, build, and operate secure, highly available Azure cloud architecture and Kubernetes clusters for microservices, event-driven systems, and AI workloads. Automate provisioning and lifecycle management using Infrastructure as Code, GitOps, and Helm while embedding security into CI/CD pipelines through scanning, policy-as-code, and deployment gates. Own the security posture across identity, networking, and supply chain controls, and lead incident response, root cause analysis, and reliability improvements using comprehensive observability tools. Define SLIs, SLOs, and operational standards to ensure safe production operations and evolve engineering culture toward reliability-first, security-by-default delivery. Partner with data/ML teams to govern emerging AI and LLM workloads, reducing risks such as prompt injection and uncontrolled compute spend.
Required Qualifications
- 5+ years in DevOps, SRE, Platform, Cloud or Infrastructure Engineering roles with significant production ownership
- Deep hands-on Azure experience, especially AKS, Azure Networking, Entra ID, Azure Policy, Azure Monitor, Application Insights, Log Analytics and Azure DevOps
- Strong AKS and Kubernetes experience, including cluster and node pool design, CNI/networking, private clusters, DNS, ingress, RBAC, Managed Identities and autoscaling
- Strong infrastructure security background on Azure: Defender for Cloud, Microsoft Sentinel or SIEM/SOAR, Zero Trust, least privilege, network hardening, secrets management and supply chain security
- Infrastructure as Code and configuration management experience with Terraform, Bicep or Pulumi, plus Helm and/or Kustomize
- CI/CD and DevSecOps experience, ideally with Azure DevOps Pipelines, including SAST/DAST, dependency or image scanning, IaC scanning, secret detection and deployment gates
- Observability and incident management experience across metrics, logs, traces, alerting, SLOs/SLIs, root cause analysis and durable remediation
- Strong scripting or programming ability in Python, Bash, PowerShell and/or Go
- Practical experience with Kubernetes API gateways or ingress platforms, such as Apache APISIX, Kong or similar
Desired Qualifications
- Experience deploying or governing AI, LLM or agentic workloads on Kubernetes, including inference serving, tool gateways or MCP servers, guardrails, sandboxing, agent identity and behavioral observability
- Relevant certifications such as CKA, CKS, AZ-104, AZ-400, AZ-500, AZ-305, SC-100 or equivalent
- Service mesh experience with Istio, Linkerd or Cilium for traffic management, observability and mTLS
- GitOps experience with Argo CD or Flux
- Experience with event-driven systems such as Kafka, Azure Event Hubs, Azure Service Bus or RabbitMQ
- Experience with multi-cluster or hybrid-cloud Kubernetes, chaos engineering, resilience testing or FinOps/cost governance on Azure
- Performance tuning experience for high-throughput APIs, event-processing platforms or Kubernetes workloads
- A calm, structured approach during high-pressure incidents, including security events
- A security-first mindset with the ability to balance reliability, delivery speed and risk
- Clear communication skills and the ability to simplify complex technical topics for different audiences
- A collaborative, ownership-driven style: you work well in a platform team, share knowledge openly and can take a problem from architecture through production operation
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.