People Data Labs logo
People Data LabsPosted 1 month ago

Senior Software Engineer, Data Acquisition

$160,000–$200,000 year

RemoteUnited States

Full TimeSenior LevelSmall

Job Summary

Contribute to the architecture and improvement of the data acquisition and processing platform, increasing reliability, throughput, and observability. Use and develop web crawling technologies to capture and catalog internet data, while building, operating, and evolving large-scale distributed systems that collect, process, and deliver web data. Design backend services managing distributed job orchestration, data pipelines, and asynchronous workloads, ensuring high quality and consistency across structured datasets. Continuously improve ingestion system speed, scalability, and fault tolerance, then partner with product and engineering teams to design new data products powered by collected data.

Required Qualifications

  • 7+ years of professional experience building or operating backend or infrastructure systems at scale
  • Solid programming experience in Python, Go, Rust, or similar, including experience with async / await, coroutines, or concurrency frameworks
  • Strong grasp of software architecture and backend fundamentals; you can reason clearly about concurrency, scalability, and fault tolerance
  • Solid understanding of browser rendering pipeline, web application architecture (auth, cookies, http request / response)
  • Familiarity with network architecture and debugging (HTTP, DNS, proxies, packet capture and analysis)
  • Solid understanding of distributed systems concepts: parallelism, asynchronous programming, backpressure, and message-driven design
  • Experience designing or maintaining resilient data ingestion, API integration, or ETL systems
  • Proficiency with Linux / Unix command-line tools and system resource management
  • Familiarity with message queues, orchestration, and distributed task systems (Kafka, SQS, Airflow, etc.)
  • Experience evaluating and monitoring data quality, ensuring consistency, completeness, and reliability across releases
  • Work independently in a fast-paced, remote-first environment, proactively unblocking themselves and collaborating asynchronously
  • Communicate clearly and thoughtfully in writing (Slack, docs, design proposals)
  • Write and maintain technical design documents, including pipeline design, schema design, and data flow diagrams
  • Scope and break down complex projects into deliverable milestones, and communicate progress, risks, and blockers effectively
  • Balance pragmatism with craftsmanship, shipping reliable systems while continuously improving them

Desired Qualifications

  • Degree in a quantitative field such as computer science, mathematics, or engineering
  • Experience as a Red Teamer
  • Experience working on large-scale data ingestion, crawling, or indexing systems
  • Experience with Apache Spark, Databricks, or other distributed data platforms
  • Experience with streaming data systems (Kafka, Pub/Sub, Spark Streaming, etc.)
  • Proficiency with SQL and data warehousing (Snowflake, Redshift, BigQuery, or similar)
  • Experience with cloud platforms (AWS preferred, GCP or Azure also great)
  • Understanding of modern data storage and design patterns (parquet, Delta Lake, partitioning, incremental updates)
  • Knowledge of modern data design and storage patterns (e.g., incremental updating, partitioning and segmentation, rebuilds and backfills)
  • Experience building and maintaining data pipelines on modern big-data or cloud platforms (Databricks, Spark, or equivalent)

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce