DEFCON AI logo
DEFCON AIPosted 2 weeks ago

Data & ML Engineer

$150,000–$200,000 year

RemoteUnited States

Full TimeSmallAI Software

Job Summary

Design and maintain the graph of entities and records, implementing probabilistic matching with blocking, scoring, and clustering to resolve low-signal data. Develop relevance models, calibrate thresholds, and establish abstention policies that route uncertain cases to human review. Build secure ingestion pipelines for structured and unstructured sources, enforcing schema validation, lineage capture, and audit logging. Implement embeddings and retrieval across a provenance-tracked evidence base, binding generated text to cited sources while treating insufficient evidence as a valid response. Own model packaging, versioning, and rollback within a controlled government cloud environment. Submit changes through a gated release process and instrument telemetry to surface degradation early.

Required Qualifications

  • 5+ years of experience in data engineering, data architecture, applied machine learning, ML engineering, or production analytics engineering
  • Strong Python and SQL, with demonstrated experience working with large, imperfect operational data
  • Experience delivering systems for sustained operational use rather than exploratory analysis alone
  • Routine use of AI-assisted development, with informed judgment about where it adds value and where its output requires verification
  • Ability to explain a technical decision to a stakeholder who must defend that decision without understanding its internals
  • US Citizenship
  • Active US Secret clearance
  • Willingness to travel up to 25% to customer sites, DEFCON AI HQ, and vendor facilities as required

Desired Qualifications

  • Clearance: active Top Secret
  • Matching: direct experience applying probabilistic matching to inconsistent identity data, including names, dates, addresses, and identifiers, and familiarity with the failure modes of each. Record linkage, master data management, or identity management. Graph data modeling. PostgreSQL and pgvector or comparable. Graph algorithms applied in production
  • Modeling: model calibration and threshold design. Cost-sensitive learning where error types carry unequal consequences. scikit-learn, XGBoost, PyTorch
  • Retrieval and generation: retrieval-augmented generation in production. Prompt and output-schema design. Establishing that generated output remains grounded in its sources, and testing to confirm it. Self-hosted or open-weight model operation. Fine-tuning, adapters, or custom embeddings
  • Pipelines: AWS Glue, Airflow, dbt, Spark, Kafka, or NiFi. Unstructured and semi-structured document ingestion. Synthetic or representative test data generation
  • Environment: federal DevSecOps, RMF, ATO, or DoW cloud environments. Hardened base images. Experience advancing a pipeline from development through accreditation and deployment
  • Domain: sensitive federal or defense data, and work performed under privacy or comparable handling constraints
  • Responsible AI: documentation, model cards, fairness testing, and model monitoring. NIST AI RMF or comparable practice

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce