Data Engineer
$140,000–$155,000 year
HybridNew York, United States
Job Summary
Design graph-based data structures encoding relationships across markets, assets, tenants, and transactions. Build retrieval pipelines (RAG, hybrid search) that give LLMs accurate, contextually rich grounding. Develop rule-based and agentic LLM workflows to automate investment analytical tasks, owning prompt engineering and production reliability. Construct ETL/ELT workflows ingesting large-scale internal and third-party datasets with clear ownership and quality SLAs. Implement systems for data trust and provenance, including lineage tracking and confidence scoring, treating bad data as production incidents. Develop and maintain low-latency APIs exposing ML outputs to front-end applications. Work directly with investment and asset management teams to identify analytical gaps and iterate quickly on solutions combining classical ML, graph reasoning, and LLM-native workflows.
Required Qualifications
- 4+ years in ML engineering, backend engineering, or a role spanning both
- Hands-on experience shipping LLM-powered applications in production — RAG pipelines, prompt engineering, eval frameworks
- Strong Python skills
- comfortable owning backend services and APIs end-to-end
- Experience with knowledge graphs or graph databases (Neo4j or similar)
- Proficiency building data pipelines at scale (Spark, Databricks, or equivalent)
- Deep sensitivity to data provenance — a track record of building systems that create analyst trust, not just claim it
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.