Sr. Monitoring and Observability Engineer
$31,200–$31,200 year
On-siteBengaluru, Karnataka, India
Job Summary
Design and maintain comprehensive monitoring strategies across Linux, storage, applications, and networking platforms using Prometheus, Grafana, Nagios, and SolarWinds. Architect Kafka-based event streaming pipelines for real-time data collection and integrate AI monitoring capabilities for intelligent alerting and anomaly detection. Define organization-wide KPI and metrics standards, establish SLAs and SLOs, and develop custom automation scripts in Bash, Python, or Go to reduce operational overhead. Collaborate with DevOps and application teams to optimize infrastructure performance and ensure observability across the entire stack. Lead strategic initiatives to improve user experience and system reliability while mentoring team members on best practices.
Required Qualifications
- Minimum 8+ years of professional experience in monitoring, observability, systems engineering, or DevOps roles
- Proven experience leading and mentoring technical teams in monitoring or infrastructure domains
- Deep hands-on expertise with opensource monitoring tools such as Prometheus, Grafana, and Nagios
- Strong experience with enterprise monitoring platforms like SolarWinds
- Proficiency in bash scripting and Linux system administration (RHEL/CentOS/Ubuntu preferred)
- Experience with event-driven architectures and Apache Kafka for real-time data processing
- Knowledge of monitoring and observability across Linux systems, storage infrastructure, applications, and network platforms
- Understanding of metrics collection, time-series data management, and data-driven decision-making
- Experience implementing AI/ML-based monitoring solutions or anomaly detection systems
- Strong problem-solving skills and ability to work effectively in fast-paced, complex environments
- Excellent communication skills with ability to present technical concepts to both technical and non-technical stakeholders
- Experience with Infrastructure-as-Code (IaC) tools (Terraform, Ansible, etc.)
- Certification in relevant areas (e.g., Kubernetes, cloud platforms, monitoring tools)
- Contribution to open-source monitoring projects
- 10% of the time travel
Desired Qualifications
- Experience with containerized environments and Kubernetes monitoring
- Proficiency in Python or Go for custom tooling development
- Experience with cloud platforms (AWS, Azure, GCP) and their native monitoring services
- Experience with observability backends such as OpenTelemetry
- Knowledge of application performance monitoring (APM) tools
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.