Cisco logo
CiscoPosted 1 month ago

Senior Site Reliability Engineer

On-siteBengaluru, Karnataka, India

Full TimeSenior LevelBachelors DegreeEnterprise

Job Summary

Own and drive AWS cost optimization and efficiency initiatives across ThousandEyes infrastructure by analyzing cloud usage across compute, storage, databases, observability, networking, and data platform workloads to identify savings opportunities. Partner with engineering teams to improve application and infrastructure performance while reducing cloud wastage, and drive FinOps practices including cost visibility, tagging hygiene, showback/chargeback, budgeting, forecasting, anomaly detection, and cost allocation. Improve infrastructure efficiency through better resource utilization, autoscaling, right-sizing, and capacity planning, while building automation, dashboards, reports, and guardrails to improve cost governance and operational visibility. Support ThousandEyes cost reviews, OKR tracking, leadership updates, and stakeholder communications, and collaborate with product, finance, engineering, and platform teams to align cost optimization with business priorities. Provide senior-level technical leadership, mentor engineers, and independently drive cross-functional initiatives to closure, balancing reliability, scalability, performance, and cost efficiency in all platform decisions.

Required Qualifications

  • Bachelor's degree or higher in Engineering, Computer Science, or equivalent practical experience
  • 8–12 years of relevant experience in Site Reliability Engineering, Cloud Infrastructure, Platform Engineering, DevOps, Production Engineering, or Performance Engineering
  • Strong experience with AWS services such as EC2, S3, RDS, EMR, Lambda, CloudWatch, OpenSearch, ElastiCache, IAM, VPC, and related cloud-native services
  • Practical experience in cloud cost optimization, FinOps, AWS billing analysis, cost allocation, tagging, budget tracking, forecasting, and cost governance
  • Experience with performance analysis, capacity planning, infrastructure optimization, and reliability improvements for large-scale cloud platforms
  • Experience with infrastructure-as-code and automation tools such as Terraform, CloudFormation, Puppet, Ansible, or similar
  • Strong scripting or programming skills in Python, Go, Shell, or similar languages for automation, reporting, and operational tooling
  • Experience with observability and monitoring platforms such as ThousandEyes, CloudWatch, Prometheus, Grafana, Splunk, OpenSearch, Datadog, or similar
  • Strong Linux systems knowledge and understanding of distributed systems
  • Experience in incident management, production support, reliability engineering, and operational excellence
  • Ability to analyze large-scale infrastructure, performance, and cost data and convert findings into actionable engineering recommendations
  • Strong communication skills with the ability to present technical, performance, and cost insights to engineering teams, finance, leadership, and multi-functional stakeholders
  • Experience working in Agile/Scrum environments and managing priorities across multiple teams

Desired Qualifications

  • Experience supporting or optimizing large-scale SaaS platforms in AWS
  • FinOps certification or equivalent experience with cloud financial management practices
  • Experience with AWS Savings Plans, Reserved Instances, Spot adoption, Graviton migration, storage lifecycle optimization, workload right-sizing, and commitment planning
  • Experience with cloud cost management tools such as AWS Cost Explorer, AWS CUR, Cloud-ability, Cloud-Health, or similar platforms
  • Experience with performance tuning of cloud infrastructure, distributed systems, databases, data pipelines, and high-scale services
  • Experience with data platforms or large-scale infrastructure components such as EMR, Kafka, OpenSearch, RDS, Airflow, Spark, Redis, Cassandra, or similar
  • Experience driving cross-team cost optimization programs, efficiency OKRs, governance reviews, executive reporting, and measurable savings outcomes
  • Ability to influence engineering teams toward cost-conscious architecture, performance-aware design, and operational standard methodologies
  • Exposure to AI/ML infrastructure cost optimization, capacity planning, GPU/accelerator cost governance, or AI-driven efficiency tooling

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce