Sr. Site Reliability Engineer - Core Platform & Embedded Reliability (Hybrid)
$140,000–$215,000 year
HybridAustin, Texas, United States
Job Summary
Partner with engineering leadership to define multi-year reliability roadmaps and shape architectural choices for shared components across the Falcon Platform. Design and implement improvements to services, libraries, and platforms while extending new tools for cross-cutting concerns like Unified Search and Protobuf. Establish foundational observability practices, define error budgets, and lead performance optimization through profiling and capacity planning. Conduct resilience engineering via chaos experiments and failure injection to eliminate entire classes of failures. Provide technical leadership during complex incidents, mentor engineers, and drive strategic decisions on infrastructure and operational improvements.
Required Qualifications
- 10+ years of experience building and operating distributed systems and service-oriented backends at scale
- 5+ years developing microservices for a SaaS product in a modern backend language (Go, Java, Scala, Kotlin, Python, Node.js)
- Expert-level proficiency in at least one programming language, with expert-level Go or demonstrated ability and willingness to reach expert level in Go
- Deep understanding of distributed systems: consensus algorithms, replication, consistency models, failure modes, and scalability patterns
- Proven experience scaling backend systems — sharding, partitioning, horizontal scaling, capacity planning, and performance optimization are second nature
- Deep understanding of multi-threading, concurrency, and parallel processing
- Track record of making impactful architectural decisions at organizational scope and seeing them through to production
- Strong systems thinking and the ability to influence without direct authority across organizational boundaries
- Thorough command of engineering best practices: appropriate testing paradigms, effective peer code review, and resilient architecture
- Ability to thrive in a fast-paced, test-driven, collaborative, and iterative environment; strong team-player orientation
- A desire to ship code and a love of seeing your bits run in production
- Degree in Computer Science, or commensurate experience in data structures, algorithms, and distributed systems
- Proven experience utilizing AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency and drive business outcomes
- Must be based in Midtown Manhattan, NY; Redmond, WA; Sunnyvale, CA; or Austin, TX
- It's a hybrid role, with employees typically working 2 to 3 days per week in the office
Desired Qualifications
- Experience driving reliability improvements in organizations with hundreds or thousands of microservices
- Deep knowledge of Kubernetes or other large-scale orchestration systems
- Hands-on experience with AWS, Cassandra, Kafka, Elasticsearch/OpenSearch, or similar large-scale distributed technologies
- Experience with Google Cloud Platform (GCP)
- Experience with Oracle Cloud Infrastructure (OCI)
- Experience delivering or operating services across multiple cloud providers, including multi-cloud abstraction layers, portability, and cloud-agnostic tooling
- Track record of building internal platforms, developer platforms, or tools that other engineers depend on
- Experience with infrastructure cost optimization at scale
- Background in performance engineering: profiling, optimization, and identifying system bottlenecks
- Experience with chaos engineering or resilience testing practices
- History of establishing SLO/SLI frameworks and error budgets in production environments
- Contributions to the open source community (GitHub, Stack Overflow, technical blogging)
- Prior experience in the cybersecurity or intelligence fields
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.