Director, AI Platform Reliability
$247,500–$275,000 year
On-siteSan Francisco, California, United States
Job Summary
Lead and scale multiple engineering teams responsible for high-volume, business-critical distributed systems and data platforms. Define technical strategy and architecture for systems processing hundreds of millions of transactions and terabytes of data, guiding development of Java-based microservices, Kafka streaming pipelines, and cloud-native services. Establish reliable data ingestion, transformation, and governance practices while building low-latency, fault-tolerant systems with strong disaster-recovery capabilities. Own operational SLAs, SLOs, and performance metrics, driving capacity planning and throughput optimization. Recruit and mentor engineering leaders, improve developer productivity through CI/CD automation, and manage technical debt and cloud costs. Partner with Product, SRE, Security, and Infrastructure teams to deliver strategic platform initiatives.
Required Qualifications
- 10+ years of professional software-engineering experience
- significant experience building large-scale distributed systems
- Experience leading engineering teams, architects, and staff engineers
- Demonstrated success delivering and operating platforms that process hundreds of millions of transactions, requests, or events
- Deep technical expertise in Java, JVM performance, concurrency, multithreading, memory management, and application profiling
- Strong experience designing and operating microservice-based and event-driven architectures
- Extensive production experience with Apache Kafka or a comparable distributed streaming platform
- Strong understanding of Kafka partitioning, replication, consumer groups, offset management, ordering, delivery semantics, schema evolution, and reprocessing
- Experience designing low-latency, highly available APIs and backend services
- Experience managing terabyte- or petabyte-scale datasets across relational, NoSQL, streaming, and object-storage technologies
- Strong understanding of distributed-systems concepts, including consensus, replication, partitioning, consistency models, idempotency, backpressure, and fault tolerance
- Experience operating cloud-native applications using Kubernetes, containers, infrastructure as code, and automated CI/CD pipelines
- Experience with at least one major cloud platform, such as AWS, Google Cloud, or Microsoft Azure
- Strong knowledge of observability practices involving metrics, logs, traces, profiling, alerting, dashboards, and service-level objectives
- Demonstrated experience improving system reliability, scalability, latency, cost efficiency, and engineering productivity
- Strong written and verbal communication skills, including the ability to explain complex technical decisions to engineering teams, executives, and business stakeholders
- Proven ability to build inclusive, accountable, and high-performing engineering organizations
- Candidates who currently hold valid U.S. work authorization that can be transferred to a new employer (such as certain H-1B statuses)
- Candidates authorized to work in the United States on a full-time, permanent basis without requiring new or initial employer-sponsored work authorization
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.