Sr Software Engineer
On-siteGurugram, Haryana, India
Job Summary
Design Kubernetes Operators using Kubebuilder or Operator SDK to manage complex, stateful application lifecycles and automate operational knowledge into code. Build high-scale SaaS products with deep proficiency in Go or JVM-based languages, focusing on memory-efficient, thread-safe systems programming for low latency and eventual consistency. Engineer deep-tier cloud storage solutions by understanding S3, NFS, and enterprise storage internals like NetApp OnTap or AWS FSx, while managing data persistence across hybrid environments. Implement expert-level observability by designing telemetry with OpenTelemetry for distributed tracing and scaling Prometheus and Grafana for high-cardinality metrics.
Required Qualifications
- BS degree in Computer Science, related technical field, or equivalent practical experience as a professional Platform Engineer
- 6-9 years of programming experience in Golang or Java
- Proven experience in building and operating a SaaS product at scale
- Proven cloud experience using observability and operating workloads in Kubernetes
- Experience with cloud storage S3, NFS, NetApp OnTap, AWS FSX
- Distributed Systems and programming techniques
- Monitoring and metrics infrastructure (OTEL, Prometheus, Grafana)
- Cloud (AWS / GCP/ Azure / Private DC)
- Custom Operators: Proven ability to write Kubernetes Operators (using Kubebuilder or Operator SDK) to manage complex, stateful application lifecycles
- Control Plane Logic: Deep understanding of reconciliation loops, custom resource definitions (CRDs), and how to automate operational knowledge into code
- High-Scale SaaS: 6+ years of experience building and operating production-grade SaaS products
- Systems Programming: Advanced proficiency in Go or JVM-based languages, with a focus on writing memory-efficient and thread-safe code
- Storage Internals: We need more than S3 API knowledge. You should understand the performance characteristics, limitations, and cost-drivers of S3, NFS, and enterprise storage like NetApp OnTap or AWS FSx
- Data Locality: Experience managing data persistence across hybrid environments (Public Cloud vs. Private DC)
- Instrumentation: Expert-level experience with OpenTelemetry. You don't just look at metrics; you design the telemetry that allows for deep-dive distributed tracing and root-cause analysis
- Observability Infrastructure: Hands-on experience scaling Prometheus and Grafana to handle high-cardinality data
Desired Qualifications
- Bonus: Experience with AI/ML, Natural Language
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.