Baselayer logo
BaselayerPosted 3 weeks ago

Senior AI Engineer, Agentic Data Enrichment

$230,000–$340,000 year

HybridSan Francisco, California, United States

Full TimeSenior LevelSmall

Job Summary

Own industry/category classification of businesses from heterogeneous signals and build discovery systems that filter aggregators and impersonators. Link individuals to businesses via public web evidence and develop risk/legitimacy scoring derived from web-presence signals fed into downstream underwriting. Build and evolve the shared agent infrastructure including provider-agnostic base agents, a toolset registry, and an eval harness for token-and-tool tracing. Select models and design agents with deep knowledge of failure modes, managing cost control and latency optimization across your enrichment surface. Ship LLM-driven agents to production with real users and on-call rotation, leveraging strong async Python and browser automation experience. Work on foundational graph AI problems at Baselayer, where ownership is real and the infrastructure you build becomes load-bearing for businesses.

Required Qualifications

  • Shipped LLM-driven agents to production - not notebooks, not demos. Real users, real cost, real failure modes, real on-call.
  • Strong async Python including structured-data libraries, modern web frameworks, and relational databases.
  • Experience across multiple frontier LLM providers and at least one agent framework, with deep knowledge of failure modes.
  • Built or maintained eval methodology: curated golden datasets, scoring functions, labelling guidelines, regression diagnostics.
  • Browser automation experience: headless browsers, anti-bot evasion, authenticated flows.
  • Holds informed opinions on structured-output reliability - when to use JSON-schema mode vs. function calling vs. extractor-on-top-of-text.
  • Web scraping at scale: anti-bot evasion, residential proxies, request fingerprinting, authenticated flows, CDN defeats.
  • Eval-framework experience (e.g., LangSmith, Braintrust, Evals, or custom).
  • Entity resolution / record linkage / fuzzy matching at scale.
  • Browser-automation experience at the devtools-protocol level.
  • Built a tool registry or toolset abstraction over multiple LLM providers.
  • Cost/latency optimization: response caching, semantic caching, model routing (cheap-first then escalate), thinking-budget tuning, prompt-cache hit-rate work.
  • Based in SF; hybrid - 4 days per week in office.

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce