Everpure logo
EverpurePosted 1 month ago
EXPIRED

Member of Technical Staff, Production & Platform Engineering Lead

On-siteBengaluru, Karnataka, India

Full TimeSenior LevelEnterpriseTechnology

Job Summary

Build internal APIs and abstractions that allow Software Engineers to provision AI-ready environments with a single command using Go and Python. Develop AI Agents that automate root-cause analysis of infrastructure failures and optimize performance. Design asynchronous AI processing pipelines leveraging Kafka or RabbitMQ for long-running document ingestion and RAG systems. Ensure "Golden Path" deployments for AI models using Docker and Kubernetes, syncing models, prompts, and code across environments. Manage GPU workloads and vector database storage while monitoring token latency and model drift via Prometheus and Grafana. This role bridges high-velocity AI experimentation with stable, enterprise-grade production infrastructure within the data storage industry.

Required Qualifications

  • 8+ years in Platform Infrastructure Engineering with a shift toward AI Application development
  • Strong proficiency in Python (for AI logic)
  • Strong proficiency in Go (for platform tools)
  • Practical experience with RAG (Retrieval-Augmented Generation)
  • Practical experience with Prompt Engineering
  • Experience integrating LLM APIs (OpenAI, Anthropic, or local models via Ollama)
  • Experience designing and implementing high-performance Go services that listen to Kafka/RabbitMQ streams to trigger dynamic infrastructure scaling based on real-time AI model demand
  • Experience managing and optimizing the infrastructure for Vector Databases and distributed caches (Redis), ensuring high availability for RAG data
  • Experience replacing brittle shell scripts with robust, type-safe Internal Tooling in Go for automated environment provisioning and disaster recovery
  • Deep experience with Kubernetes, specifically managing GPU workloads and specialized storage for vector databases
  • Proven experience with Kafka or RabbitMQ for managing high-volume data streams
  • Experience using Prometheus and Grafana to monitor not just system health, but AI-specific metrics like token latency and model 'drift'
  • Ability to write clean, maintainable, and testable Go code to manage complex cloud environments
  • Understanding of how to use Go's goroutines and channels to handle thousands of concurrent events
  • Builder's perspective to Kubernetes with ability to extend the K8s API with custom controllers to make the cluster 'AI-aware'
  • Understanding of the unique infrastructure needs of AI (GPU memory management, high-speed NVMe storage, and vector retrieval)
  • Background in Monitoring (Grafana/Prometheus) and Log Management (ELK)
  • Understanding of asynchronous architecture, knowing exactly when to use Kafka for high-volume streaming versus RabbitMQ for complex task routing
  • 'Security-as-Code' approach, ensuring that data privacy (vital for AI) is baked into the infrastructure layer rather than bolted on at the end
  • Mindset of collaboration, reliability, and continuous improvement to strengthen team productivity and delivery speed
  • #LI-ONSITE
  • #LI-KT7

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Find similar roles