Senior Software Engineer II - Streaming and Durability
$176,000–$196,100 year
On-siteManhattan, New York, United States
Job Summary
Design and operate the database and caching fleet (RDS/Aurora PostgreSQL, ElastiCache Redis/ValKey, DynamoDB) that underpins Compass's core products, making decisions about replication, failover, capacity, and engine strategy. Build multi-tenant data isolation into the data layer from the ground up, shipping SDKs in Go, Java, and Python that enforce tenant-safe access patterns. Own production database upgrades at fleet scale, coordinating engine version bumps and Terraform state mutations across hundreds of clusters with zero customer-facing impact. Build and scale Kafka/MSK streaming infrastructure, including topic architecture, consumer patterns, and blue/green deployment strategies. Develop deep observability into the data layer, including SLI dashboards, tenant-dimensioned telemetry, and real-time performance insights. Partner directly with product and platform teams to assess database readiness for scaling events, migrate workloads onto golden-path infrastructure, and resolve the hardest production incidents in the stack.
Required Qualifications
- 5+ years of experience building and operating distributed systems
- depth in at least one major managed database (PostgreSQL or similar) at production scale in AWS
- You've built shared tooling (SDKs, client libraries, platform APIs)
- understanding of what it takes to make infrastructure accessible and safe for engineers who aren't infrastructure specialists
- ability to think in trade-offs, not absolutes
- ability to articulate why you'd choose a soft application-enforced boundary over a hard infrastructure lock and when that calculus changes
- You write Go (or can ramp quickly)
- comfort in Terraform
- comfort in Kubernetes
- comfort in IAM-heavy AWS environments
- experience leading cross-team operational programs
- experience with change management
- experience with migration coordination
- experience with rollback planning
Desired Qualifications
- You've worked on multi-tenant isolation problems such as key prefixing, row-level security, and RBAC/ACL models
- opinions about where application-layer enforcement breaks down
- You've operated Kafka/MSK at scale
- experience dealing with consumer lag
- experience with partition rebalancing
- experience with cross-account IAM auth
- You've built or contributed to observability platforms
- specifically for database and cache monitoring
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.