Senior Software Engineer - AI Agent Platform
On-siteTel Aviv, Tel Aviv, Israel
Job Summary
Own and evolve the production runtime responsible for executing and orchestrating AI agents, including tool calling, streaming, state management, and long-running workflows. Build resilient integrations with multiple LLM providers, implementing retries, timeouts, circuit breakers, and fallback strategies based on availability, latency, and cost. Create end-to-end observability for AI requests, defining dashboards, alerts, SLOs, and runbooks for production workloads. Lead investigation of complex production issues across application code, infrastructure, and agent behavior, turning incidents into architectural improvements. Mentor engineers and establish best practices for building reliable production AI systems while collaborating with product and infrastructure teams.
Required Qualifications
- 7+ years of professional software engineering experience, primarily in backend, platform, or distributed systems
- Expert-level TypeScript and Node.js skills
- Experience with NestJS or a comparable backend framework
- Proven experience designing, building, and operating large production services
- Strong understanding of distributed-systems patterns, including retries, backoff, idempotency, circuit breakers, caching, consistency, and failure recovery
- A systematic debugging mindset and the ability to trace failures across multiple services and dependencies
- Experience owning customer-facing systems where availability, latency, and correctness directly affect users
- Strong experience with cloud infrastructure and managed services, preferably AWS
- Experience with distributed caching and storage technologies such as Redis and S3
- Hands-on experience with production observability: structured logs, metrics, tracing, dashboards, alerts, and SLOs
- Experience participating in incident response and driving follow-up improvements
- Strong API design, testing, and software architecture fundamentals
- Excellent communication and collaboration skills across engineering, product, infrastructure, and AI teams
- Bachelor's degree in Computer Science or a related field, or equivalent practical experience
- LLM APIs, streaming, tool calling, and structured outputs
- Agentic patterns such as reasoning loops, tool orchestration, and multi-step workflows
- Tokens, context windows, rate limits, latency, and model-specific behavior
- Multi-provider LLM integrations and the trade-offs between providers and models
- How retries and fallbacks can affect response quality, latency, correctness, and cost
- The observability required to understand a request across multiple LLM and tool calls
- Techniques for controlling and optimizing LLM usage and cost
Desired Qualifications
- Experience building an AI gateway, agent runtime, inference platform, or workflow engine
- Experience with OpenAI, Gemini/Vertex AI, AWS Bedrock, or similar platforms
- Familiarity with model and provider routing based on quality, availability, latency, and cost
- Experience with Kafka or other event-driven architectures
- Familiarity with MCP or agent-to-agent communication protocols
- Experience with Kubernetes and cloud-native infrastructure
- Experience with Grafana, or similar tooling
- Experience with chaos engineering, fault injection, or large-scale load testing
- Experience building internal developer platforms or frameworks used by multiple engineering teams
- Familiarity with conversational AI, RAG, memory systems, or generative user experiences
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.