LogicMonitor logo
LogicMonitorPosted 1 week ago

Director, AI Platform Reliability

$247,500–$275,000 year

On-siteSan Francisco, California, United States

Full TimeSenior LevelLarge

Job Summary

Lead and scale multiple engineering teams responsible for high-volume, business-critical distributed systems and data platforms. Define technical strategy and architecture for systems processing hundreds of millions of transactions and terabytes of data, guiding development of Java-based microservices, Kafka streaming pipelines, and cloud-native services. Establish reliable data ingestion, transformation, and governance practices while building low-latency, fault-tolerant systems with strong disaster-recovery capabilities. Own operational SLAs, SLOs, and performance metrics, driving capacity planning and throughput optimization. Recruit and mentor engineering leaders, improve developer productivity through CI/CD automation, and manage technical debt and cloud costs. Partner with Product, SRE, Security, and Infrastructure teams to deliver strategic platform initiatives.

Required Qualifications

  • 10+ years of professional software-engineering experience
  • significant experience building large-scale distributed systems
  • Experience leading engineering teams, architects, and staff engineers
  • Demonstrated success delivering and operating platforms that process hundreds of millions of transactions, requests, or events
  • Deep technical expertise in Java, JVM performance, concurrency, multithreading, memory management, and application profiling
  • Strong experience designing and operating microservice-based and event-driven architectures
  • Extensive production experience with Apache Kafka or a comparable distributed streaming platform
  • Strong understanding of Kafka partitioning, replication, consumer groups, offset management, ordering, delivery semantics, schema evolution, and reprocessing
  • Experience designing low-latency, highly available APIs and backend services
  • Experience managing terabyte- or petabyte-scale datasets across relational, NoSQL, streaming, and object-storage technologies
  • Strong understanding of distributed-systems concepts, including consensus, replication, partitioning, consistency models, idempotency, backpressure, and fault tolerance
  • Experience operating cloud-native applications using Kubernetes, containers, infrastructure as code, and automated CI/CD pipelines
  • Experience with at least one major cloud platform, such as AWS, Google Cloud, or Microsoft Azure
  • Strong knowledge of observability practices involving metrics, logs, traces, profiling, alerting, dashboards, and service-level objectives
  • Demonstrated experience improving system reliability, scalability, latency, cost efficiency, and engineering productivity
  • Strong written and verbal communication skills, including the ability to explain complex technical decisions to engineering teams, executives, and business stakeholders
  • Proven ability to build inclusive, accountable, and high-performing engineering organizations
  • Candidates who currently hold valid U.S. work authorization that can be transferred to a new employer (such as certain H-1B statuses)
  • Candidates authorized to work in the United States on a full-time, permanent basis without requiring new or initial employer-sponsored work authorization

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce