Rubiscape logo
RubiscapePosted 2 months ago

MLOps Engineer

On-sitePune, Maharashtra, India

Full TimeSmall

Job Summary

Design end-to-end ML pipelines using MLflow, Kubeflow, or Airflow to handle training, validation, packaging, and deployment at scale. Build the model registry architecture within RubiStudio, defining versioning strategies, stage transitions, approval gates, and rollback mechanisms. Implement automated monitoring for data drift and prediction quality degradation, surfacing alerts into operational dashboards. Manage containerised model serving infrastructure across multi-cloud and on-premises topologies while enforcing MLOps best practices for reproducible experiments and audit-ready lineage. Collaborate with security teams to ensure compliance with enterprise data governance standards and instrument inference endpoints with SLOs for latency and throughput. Own on-call response for production model degradation incidents.

Required Qualifications

  • 3+ years in MLOps, ML infrastructure, or ML platform engineering roles with demonstrable production deployments
  • Proficiency with MLflow (or similar experiment tracking + registry tools) and workflow orchestration frameworks such as Airflow, Kubeflow Pipelines, or Prefect
  • Strong container and Kubernetes skills: writing Helm charts, managing model-serving deployments, horizontal pod autoscaling for inference workloads
  • Experience with at least one model-serving framework: TorchServe, Triton Inference Server, BentoML, or Seldon Core
  • Working knowledge of Python and shell scripting sufficient to own pipeline code, not just configure GUI tools
  • Familiarity with observability tooling (Prometheus, Grafana, OpenTelemetry) applied to ML workloads

Desired Qualifications

  • Experience operating ML infrastructure in air-gap or on-premises environments for government or defence customers
  • Knowledge of feature stores (Feast, Tecton, or a custom implementation) and their integration into training and online inference paths
  • Exposure to GPU cluster management and optimising inference throughput for large model serving
  • Certification in AWS Machine Learning Specialty, Google Professional ML Engineer, or equivalent

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce