Senior AI Engineer, Agentic Data Enrichment
$230,000–$340,000 year
HybridSan Francisco, California, United States
Job Summary
Own industry/category classification of businesses from heterogeneous signals and build discovery systems that filter aggregators and impersonators. Link individuals to businesses via public web evidence and develop risk/legitimacy scoring derived from web-presence signals fed into downstream underwriting. Build and evolve the shared agent infrastructure including provider-agnostic base agents, a toolset registry, and an eval harness for token-and-tool tracing. Select models and design agents with deep knowledge of failure modes, managing cost control and latency optimization across your enrichment surface. Ship LLM-driven agents to production with real users and on-call rotation, leveraging strong async Python and browser automation experience. Work on foundational graph AI problems at Baselayer, where ownership is real and the infrastructure you build becomes load-bearing for businesses.
Required Qualifications
- Shipped LLM-driven agents to production - not notebooks, not demos. Real users, real cost, real failure modes, real on-call.
- Strong async Python including structured-data libraries, modern web frameworks, and relational databases.
- Experience across multiple frontier LLM providers and at least one agent framework, with deep knowledge of failure modes.
- Built or maintained eval methodology: curated golden datasets, scoring functions, labelling guidelines, regression diagnostics.
- Browser automation experience: headless browsers, anti-bot evasion, authenticated flows.
- Holds informed opinions on structured-output reliability - when to use JSON-schema mode vs. function calling vs. extractor-on-top-of-text.
- Web scraping at scale: anti-bot evasion, residential proxies, request fingerprinting, authenticated flows, CDN defeats.
- Eval-framework experience (e.g., LangSmith, Braintrust, Evals, or custom).
- Entity resolution / record linkage / fuzzy matching at scale.
- Browser-automation experience at the devtools-protocol level.
- Built a tool registry or toolset abstraction over multiple LLM providers.
- Cost/latency optimization: response caching, semantic caching, model routing (cheap-first then escalate), thinking-budget tuning, prompt-cache hit-rate work.
- Based in SF; hybrid - 4 days per week in office.
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.