Smarsh logo
SmarshPosted 1 week ago

Senior Software Engineer, Python + AI Platform

RemoteUnited States

Full TimeSenior LevelLarge

Job Summary

Drive backend development for AI workflows by building and evolving Python/FastAPI services powering core agentic workflows and platform capabilities. Productionize LLM integrations with Bedrock usage, quotas, retries, failover, and cost controls while designing for security, compliance, and tenant isolation. Build for scale through async job orchestration, performance tuning, and data-layer optimization to handle petabyte-scale enterprise data. Support multi-tenant architecture with role-based access and SSO integration, and improve platform reliability via monitoring, tracing, and alerting for LLM pipelines. Design typed API contracts and implement real-time event delivery patterns to support live workflow state and agent feedback loops. Contribute to technical decisions on shared services and integration boundaries while translating evolving product requirements into practical solutions with stakeholders. Champion code quality through strong typing, automated testing, and continuous integration practices.

Required Qualifications

  • 7+ years professional software development
  • 5+ years building Python services in production
  • Deep experience with APIs
  • async processing
  • background jobs
  • workflow orchestration
  • AWS experience
  • Proven ability to productionize complex backend systems
  • reliability
  • observability
  • retries
  • throughput
  • failure handling
  • performance tuning
  • Strong knowledge of PostgreSQL
  • large-scale data processing patterns
  • indexing
  • query tuning
  • batch/stream tradeoffs
  • Experience with retrieval-augmented generation (RAG)
  • vector search
  • embedding-based systems
  • Experience with multi-tenant systems
  • RBAC
  • audit logging
  • secure data handling
  • regulated environments
  • Ability to work from partial requirements
  • Hands-on experience building LLM-driven workflows
  • tool-calling
  • state machines
  • human-in-the-loop approval patterns
  • checkpoint/resume
  • multi-step agent orchestration
  • Familiarity with frameworks like LangGraph or equivalent
  • Experience working on or alongside AI-native engineering teams
  • hands-on prompt engineering
  • eval design
  • LLM cost optimization
  • caching strategies
  • token efficiency
  • model selection tradeoffs

Desired Qualifications

  • LLM / AI platform experience
  • Bedrock
  • OpenAI
  • Anthropic
  • LangChain/LangGraph
  • prompt workflows
  • evals
  • tool-calling systems
  • Experience integrating external AI services safely and reliably
  • Identity and access
  • SSO/SAML/OIDC
  • enterprise auth patterns
  • Graph-shaped data and entity resolution
  • Experience with graph-backed data models
  • entity deduplication
  • mention linking
  • building systems that reason over connected, structured records
  • Observability stack
  • OpenTelemetry
  • tracing
  • metrics
  • alerting
  • cost/usage dashboards
  • Regulated communications or compliance domain
  • Background in systems that handle sensitive communications
  • audit trails
  • data subject to legal or regulatory review
  • Infrastructure as code
  • Terraform
  • feature flags
  • canary deployments
  • release strategies

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce