OKX logo
OKXPosted 1 month ago

Senior AI Platform & Agentic Infrastructure Engineer

$178,000–$321,000 year

On-siteSan Jose, California, United States

Part TimeSenior LevelLargeFintech Platform

Job Summary

Inherit, operate, and progressively migrate the working prototype estate (Google Workspace–native automation across Apps Script, Drive, and the Docs/Sheets/Slides APIs; locally scheduled jobs; and Claude Code agent tooling) to the target platform without interrupting daily and board-cycle workflows. Re-architect Hive Mind into a resilient AWS or GCP platform with high availability (HA), disaster recovery (DR), defined service-level objectives (SLOs), and full observability, and own it in production. Build the agentic runtime and harness: orchestration, multi-model routing that sends each task to the right model, tool, MCP, and plugin integration across model providers, and the evaluation, red-team, and regression harnesses that grade agents before and after production, with evaluation gates, judge calibration, and cost controls that hold for each model, including prompt-injection and data-exfiltration threat modeling for agents that read untrusted content. Build the Responsible-AI and model-governance layer: hallucination, bias, and drift controls, output validation, guardrails, and complete logging. Extend the existing governance design (human-calibrated judge gates, golden-set regression, provenance registry) rather than replacing it. New model providers enter through the same governance and approved-tooling review, not around it. Build data infrastructure with provenance and lineage: immutable audit trails, versioned evidence, reproducible pipelines, and traceability from source to report. Engineer data protection: encryption, key and secrets management, least-privilege access, sensitive-data handling, residency, and defensible retention. Stand up the computer-assisted audit technique (CAAT) and continuous-monitoring data foundation: analytics over full populations, with exceptions streamed in real time. Auditors and the hiring manager define the audit logic; you make it run at production grade. Own the cloud foundation: infrastructure-as-code, CI/CD, identity, networking, observability, and cost controls, and make the platform

Required Qualifications

  • 7+ years building and operating resilient backend or platform systems in production, including on-call ownership over time
  • Proven brownfield migrations: you have taken a founder-built or prototype system to production grade while it stayed in daily use
  • Strong engineering fundamentals: data structures and algorithms, fluent Python, strong SQL and data modeling, plus one additional systems language (TypeScript/Node or Go), with the full-stack range to connect the infrastructure yourself
  • Agentic runtime and harness engineering, proven by building: the field is too young to demand years of it, so we weigh real systems shipped over tenure. You have genuinely built model-agnostic agent orchestration, model routing and evaluation across providers, agent SDKs (Claude Agent SDK or equivalent), MCP servers, skill- and hook-based agent tooling, and the evaluation and red-team harnesses that grade agent behavior, with responsible-AI controls (hallucination, bias, drift)
  • Data engineering with provenance and lineage, and security and data-protection engineering by default (encryption, identity and access management, secrets, retention). You build systems that could withstand external audit, by design
  • Cloud (AWS or GCP) plus resilience engineering: infrastructure-as-code (Terraform), CI/CD, HA, DR, SLOs, and observability, backed by automated testing and documentation
  • You ship inside locked-down enterprise environments (TLS-intercepting proxies, endpoint detection and response tooling, restricted installs, security guardrails, OAuth admin consent) without treating security as someone else's problem
  • Partnership: you translate audit needs into systems, explain technical risk to non-engineers, and drive adoption
  • Experience passing external audit, SOC 2, or SOX (helpful, not required)

Desired Qualifications

  • Multi-agent orchestration frameworks, MCP servers, and tooling and plugin development across the Anthropic, OpenAI, and Google model ecosystems
  • Model-risk or AI-governance program experience; frameworks such as the NIST AI Risk Management Framework
  • Continuous-auditing or continuous-controls-monitoring platforms; streaming and real-time data at scale; statistical anomaly detection and applied machine learning beyond LLMs; vector stores and retrieval; LLM cost engineering
  • Crypto and blockchain literacy; regulated financial-services, fintech, or crypto experience

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce