Hadrian Automation logo
Hadrian AutomationPosted 3 weeks ago

ML Platform Engineer

$170,000–$300,000 year

On-siteLos Angeles, California, United States

Full TimeSmall

Job Summary

Build the production ML platform enabling Hadrian's factories to safely depend on models for drawing extraction, cycle-time prediction, forecasting, and scheduling with measurable performance and fast rollback. Develop shared batch and online serving for tabular, vision, document-AI, scheduling, graph, and embedding workloads targeting clear SLAs for latency, availability, and isolation. Create repeatable release and evaluation processes featuring automated tests, reproducible artifacts, lineage, shadow deployments, canaries, and A/B tests. Own online feature serving and maintain contract integrity with offline feature tables, proactively detecting training-serving skew, feature drift, bad data, and model degradation. Build operational tooling for telemetry, incident response, autoscaling, resource and GPU management, cost attribution, and secure model routing. Develop APIs, SDKs, reusable templates, and documentation that teams can adopt without requiring close support.

Required Qualifications

  • Track record building and operating production ML infrastructure across multiple models or inference workloads
  • Strong production-level Python and SQL skills, including typing, testing, packaging, API design, and building observability features
  • Hands-on experience with Kubernetes, containers, and handling distributed-system failure modes such as retries, partial failures, idempotence, and resource isolation
  • Engineering background with model registries, feature systems, batch/real-time inference, experiment tracking, or model CI/CD workflows
  • Practical judgment around latency, throughput, availability, multi-tenancy, autoscaling, and infrastructure cost optimizations
  • Ability to build stable interfaces and collaborate closely with engineering and scientific stakeholders
  • U.S. citizen, lawful permanent resident of the U.S., protected individual as defined by 8 U.S.C. 1324b(a)(3), or eligible to obtain the required authorizations from the U.S. Department of State

Desired Qualifications

  • Experience implementing feature stores (Feast, Tecton, or internal systems)
  • Production work with Ray Serve, KServe, Triton, BentoML, SageMaker, Vertex AI, or custom gRPC inference services
  • Experience serving and evaluating vision, document-understanding, embedding, or generative pipelines
  • Expertise in GPU inference optimization, multi-model serving, edge inference, or Go/Rust performance-sensitive AI services
  • Background in regulated environments or open-source contributions to ML infrastructure projects (MLflow, Feast, KServe, Ray)

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce