Staff / Senior Software Engineer (Agentic Search) - Crawler
On-siteAmsterdam, North Holland, The Netherlands
Job Summary
Design and operate web-scale crawling systems to acquire content from the open web and large-scale data sources. Build ingestion workflows for internal feeds, external crawlers, and partner integrations while developing crawl scheduling, prioritization, and recrawl policies. Construct systems for URL discovery, deduplication, content extraction, and crawl orchestration to manage high-throughput data flows across billions of URLs. Ensure reliable infrastructure operation under high-throughput conditions and define observability metrics for coverage, freshness, and content quality. Monitor resource usage, bandwidth, and infrastructure costs. Collaborate with indexing and ML teams to ensure acquired content meets retrieval requirements and enable safe experimentation with acquisition strategies.
Required Qualifications
- 5+ years of experience building backend or distributed systems
- Strong Go or C++ expertise
- Experience with large-scale distributed systems (10k+ RPS, billions of URLs, high-throughput pipelines)
- Understanding of web protocols (HTTP, DNS, TLS), crawling, scraping, and content extraction
- Experience operating production systems and debugging failures in distributed environments
- Strong understanding of scalability, fault tolerance, and resource management
- Applicants must be authorized to work in the country in which they apply
Desired Qualifications
- Web crawling
- Building streaming data pipelines and event-driven systems
- Kafka, Pulsar, NATS, RabbitMQ, or similar messaging platforms
- Designing distributed schedulers, queues, and asynchronous processing systems
- Spark, Flink, Beam, or MapReduce
- Ad tech, social networks, search engines, or other large-scale content platforms
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.