Medallia logo
MedalliaPosted 2 weeks ago

Senior Platform Software Engineer (Cache, Query and Compute Platforms)

$138,000–$205,000 year

HybridMcLean, Virginia, United States

Full TimeSenior LevelLarge

Job Summary

Operate and maintain Redis, Kvrocks, Trino, and Spark platforms, managing cluster deployments, replication, failover, and performance tuning. Lead production incident triage, root-cause analysis, and capacity planning while building operational automation, monitoring, and troubleshooting playbooks. Review application architectures for production readiness and validate upgrade and recovery procedures. Participate in a periodic on-call rotation to ensure 24/7 reliability of these critical services. This role is hybrid with three days per week onsite in Tysons, Virginia.

Required Qualifications

  • 5 years of experience in software, systems, platform, DevOps, or reliability engineering
  • 3 or more years of experience operating distributed platforms in production
  • Hands-on operational experience managing and scaling at least two of the following platforms in production: Redis, Kvrocks, Trino, or Spark
  • Experience leading production incident triage, root-cause analysis (RCA), and performance optimization for memory fragmentation, query execution bottlenecks, cluster failovers, and CPU/memory resource utilization
  • Experience designing, testing and executing disaster recovery plans, automated node failovers, cluster re-sharding, and zero-downtime upgrades for high-throughput platforms
  • Demonstrated experience in Java, Go, Python, or a similar programming language
  • Experience conducting formal architecture reviews, authoring technical design documents (e.g., RFCs/ADRs), and establishing operational runbooks across engineering teams
  • Candidates based in the Tysons vicinity will be prioritized as this role is Hybrid, 3 days per week onsite

Desired Qualifications

  • Experience with additional platforms among Redis, Kvrocks, Trino, and Spark
  • Experience operating distributed platforms on Kubernetes or cloud infrastructure
  • Experience integrating Trino with object storage, Hive Metastore, or external data catalogs
  • Experience validating topology, failover, recovery, and application-health behavior
  • Experience with platform automation, observability, upgrades, capacity planning, and failure-mode testing
  • Experience supporting large-scale, business-critical cache, query, or compute platforms

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce