XenonStack logo
XenonStackPosted 5 months ago

Agentic Infrastructure Observability Engineer

On-siteMohali, Punjab, India

Full TimeSmall

Job Summary

Design and implement end-to-end observability frameworks covering metrics, logs, traces, and cost telemetry for agentic systems. Build dashboards and alerting systems to monitor reliability, performance, and drift in real-time, while tracking LLM usage, context windows, and token allocation within LangChain, LangGraph, and RAG pipelines. Define SLOs, SLIs, and SLAs for agentic workflows, conduct root cause analysis of agent failures, and integrate observability into CI/CD and AgentOps pipelines. Develop custom plugins to extend monitoring for LLMs and data pipelines, collaborating with AgentOps, DevOps, and Data Engineering teams to ensure system-wide transparency. Provide executive-level reporting on reliability and efficiency metrics, implementing feedback loops to reduce downtime and stay updated with state-of-the-art AI monitoring frameworks.

Required Qualifications

  • 3–6 years of experience in SRE, DevOps, or Observability Engineering
  • Strong knowledge of observability tools (Prometheus, Grafana, ELK, OpenTelemetry, Jaeger)
  • Experience with cloud-native infrastructure (AWS, GCP, Azure) and Kubernetes monitoring
  • Proficiency in Python, Go, or Bash for scripting and automation
  • Understanding of AI/LLM pipelines, RAG systems, and vector databases
  • Hands-on with CI/CD pipelines and monitoring-as-code

Desired Qualifications

  • Experience with AgentOps tools (LangSmith, PromptLayer, Arize AI, Weights & Biases)
  • Exposure to AI-specific observability (token usage, model latency, hallucination tracking)
  • Knowledge of Responsible AI monitoring frameworks
  • Background in BFSI, GRC, SOC, or other regulated industries

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce