Senior Platform Software Engineer (Cache, Query and Compute Platforms)
$138,000–$205,000 year
HybridMcLean, Virginia, United States
Job Summary
Operate and maintain Redis, Kvrocks, Trino, and Spark platforms, managing cluster deployments, replication, failover, and performance tuning. Lead production incident triage, root-cause analysis, and capacity planning while building operational automation, monitoring, and troubleshooting playbooks. Review application architectures for production readiness and validate upgrade and recovery procedures. Participate in a periodic on-call rotation to ensure 24/7 reliability of these critical services. This role is hybrid with three days per week onsite in Tysons, Virginia.
Required Qualifications
- 5 years of experience in software, systems, platform, DevOps, or reliability engineering
- 3 or more years of experience operating distributed platforms in production
- Hands-on operational experience managing and scaling at least two of the following platforms in production: Redis, Kvrocks, Trino, or Spark
- Experience leading production incident triage, root-cause analysis (RCA), and performance optimization for memory fragmentation, query execution bottlenecks, cluster failovers, and CPU/memory resource utilization
- Experience designing, testing and executing disaster recovery plans, automated node failovers, cluster re-sharding, and zero-downtime upgrades for high-throughput platforms
- Demonstrated experience in Java, Go, Python, or a similar programming language
- Experience conducting formal architecture reviews, authoring technical design documents (e.g., RFCs/ADRs), and establishing operational runbooks across engineering teams
- Candidates based in the Tysons vicinity will be prioritized as this role is Hybrid, 3 days per week onsite
Desired Qualifications
- Experience with additional platforms among Redis, Kvrocks, Trino, and Spark
- Experience operating distributed platforms on Kubernetes or cloud infrastructure
- Experience integrating Trino with object storage, Hive Metastore, or external data catalogs
- Experience validating topology, failover, recovery, and application-health behavior
- Experience with platform automation, observability, upgrades, capacity planning, and failure-mode testing
- Experience supporting large-scale, business-critical cache, query, or compute platforms
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.