Applied Machine Learning Engineer
On-siteResearch Triangle Park, North Carolina, United States
Job Summary
Design and own Vulcan's data architecture from operational data stores through ETL pipelines to the analytics and AI layer, evaluating platforms for the data Lakehouse, ETL tooling, and operational databases while weighing scalability, compliance requirements, operational burden, and cost. Build reliable ingest paths for structured data, time-series data, files, images, and other outputs from manufacturing and lab systems, replacing manual workflows with monitored, reliable pipelines that surface quality issues early. Define data models supporting operational queries, analytical workloads, and future AI and ML applications, ensuring every data point carries necessary metadata. Collaborate across engineering, operations, and IT to translate operational requirements into durable, scalable data architecture, applying sound engineering practices including version control, testing, observability, and documentation. This role begins in Durham, NC and is expected to move to Benson, NC upon completion of the new facility.
Required Qualifications
- 8+ years of experience in data engineering, data infrastructure, or a closely related technical role with a track record of owning and delivering production systems
- Demonstrated experience designing and building data lakes, Lakehouses, or analytical data stores; understands the tradeoffs between platforms and can make and defend platform selection decisions
- Strong experience designing and building ETL/ELT pipelines that enrich and contextualize data
- Deep fluency with data modeling for both operational and analytical workloads; can design schemas that serve present needs without foreclosing future ones
- Experience with relational databases (PostgreSQL, SQL Server, or similar); writes and debugs SQL confidently
- Comfortable working in a fast-moving environment with a small team, making decisions with incomplete information and documenting them clearly for future colleagues
- Strong communicator who can work across technical and non-technical stakeholders and translate between operational requirements and data architecture decisions
- Must be a U.S. Person due to required access to U.S. export-controlled information or facilities
Desired Qualifications
- Experience with time-series databases (InfluxDB, TimescaleDB, or similar) common in industrial and IoT environments
- Familiarity with industrial data concepts — historian data, process tags, OT/IT integration — and the data challenges specific to manufacturing environments
- Experience working on or alongside a Unified Namespace or MQTT-based data architecture; understands how industrial messaging infrastructure relates to the data layer
- Familiarity with data Lakehouse platforms and open table formats (Delta Lake, Apache Iceberg, or similar)
- Experience with ETL orchestration tooling (Airflow, Prefect, dbt, or similar)
- Comfort with scripting and lightweight development (Python, SQL, or similar) for pipeline development and data quality tooling
- Familiarity with cloud platforms (AWS, Azure, or GCP) and experience evaluating on-premises vs. cloud tradeoffs for data infrastructure
- Experience working in a controlled information environment; familiarity with the handling requirements for Controlled Unclassified Information (CUI) or export-controlled technical data under ITAR or EAR
- Experience in a manufacturing, industrial, or operations-heavy environment
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.