Lead Infrastructure Engineer
On-siteAtlanta, Georgia, United States or Charlotte, North Carolina, United States
Job Summary
Design and implement enterprise infrastructure technology platforms across cloud, network, database, and middleware domains, integrating control plane components and policy-as-code enforcement engines. Develop advanced automation and monitoring techniques to ensure high availability, while leading large, complex technical initiatives involving cross-departmental teams. Troubleshoot critical infrastructure issues, collaborate with external partners to drive continuous improvement, and provide technical guidance to lower-level professionals. Document designs and communicate effectively with stakeholders to align solutions with business objectives. This role requires a minimum of 10 years of experience in infrastructure engineering and proficiency in Python, Java, Kubernetes, and OpenTelemetry.
Required Qualifications
- Bachelor's degree in Computer Science, Engineering, Information Systems, or related field
- Minimum of 10 years of professional experience in infrastructure engineering
- Advanced knowledge of enterprise infrastructure technologies including cloud, network, database, storage, platform, computing, and middleware
- English (Required)
- 1st shift (United States of America)
Desired Qualifications
- Experience in the financial services industry, particularly within regulated environments subject to FFIEC, SOX, NIST 800-53, or PCI compliance frameworks
- Strong proficiency in Python (FastAPI) and/or Java for implementing production microservices deployed on Kubernetes or OpenShift
- Hands-on implementation experience with Open Policy Agent (OPA) and Rego for policy-as-code authoring, testing, and enforcement
- Proven experience implementing Apache Kafka producer/consumer services, schema registry integration (Confluent), and event-driven architecture patterns in production environments
- Demonstrated experience implementing resilience patterns including circuit breakers, bulkheads, retry with backoff, dead-letter queues, and graceful degradation in distributed systems
- Experience implementing horizontal scaling, connection pooling, caching strategies, and performance optimization to meet sub-200ms p95 latency and 99.9%+ availability SLOs
- Proficiency implementing GitLab CI/CD pipelines, including custom pipeline stages, security scanning integration, and governance gate development
- Experience implementing Backstage developer portal workflows, including software template development, catalog entity management, and scaffolder plugin authoring
- Working knowledge of HashiCorp Vault for implementing secrets injection, credential rotation, and automated secrets path bootstrapping
- Experience implementing observability instrumentation using OpenTelemetry, Prometheus, Splunk, and distributed tracing systems (Jaeger/Tempo)
- Familiarity with cloud platforms (AWS or Azure) and infrastructure-as-code tools such as Terraform and Ansible
- Experience implementing container security practices including image signing (Cosign), SBOM generation (CycloneDX), and container scanning pipelines
- Industry certifications such as CKA (Certified Kubernetes Administrator), AWS Solutions Architect, Azure Solutions Architect, HashiCorp Certified Vault Associate, or equivalent
- Prior experience implementing or contributing to internal developer platforms (IDPs) or platform engineering initiatives at enterprise scale
- Experience engineering with AI coding agents and assistants (e.g., GitLab Duo, GitHub Copilot, or similar) as collaborative development partners to accelerate implementation, generate code, review designs, and augment engineering workflows
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.