Vialto logo
VialtoPosted 2 weeks ago

Ops / SRE Support Engineer – VLabs - Manager

On-siteBengaluru, Karnataka, India

Full TimeLarge

Job Summary

Monitor and support production systems across applications, infrastructure, and cloud services in Azure environments, responding to incidents and driving timely resolution through root-cause analysis and post-incident reviews. Maintain operational runbooks and standard operating procedures while working with logs, metrics, traces, and alerts to ensure system visibility using tools like Tempo, Loki, Prometheus, and Grafana. Participate in on-call rotations to support triage, escalation, and communication during incidents, contributing to improved MTTR and MTTD. Develop automation scripts in Python or Bash to streamline operational tasks and assist in automating incident response and remediation workflows. Operate containerized environments using Docker and Kubernetes while leveraging AI-assisted tools for incident investigation and anomaly detection.

Required Qualifications

  • 3–6+ years of experience in operations, production support, SRE, or DevOps roles
  • Hands-on experience supporting production systems in cloud environments (Azure preferred)
  • Experience with monitoring and observability tools (logs, metrics, traces, alerting)
  • Strong understanding of incident management and troubleshooting practices
  • Experience with scripting or automation (Python, Bash, etc.)
  • Familiarity with Docker and Kubernetes
  • Experience working with logs, metrics, traces, and alerts
  • Experience in incident response, triage, and root-cause analysis
  • Experience with monitoring dashboards and alerting systems
  • Understanding of SRE principles (SLIs, SLOs, error budgets)
  • Ability to troubleshoot distributed systems and service dependencies

Desired Qualifications

  • Experience with: Azure Monitor, Application Insights, Prometheus, Grafana OpenTelemetry (OTEL) instrumentation ,CI/CD tools such as Azure DevOps or GitHub Actions
  • Familiarity with: AIOps concepts (alert correlation, anomaly detection, automation)
  • Query languages like KQL or SQL
  • Experience working in high-scale or enterprise production environments

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce