Senior Decision Intelligence Engineer (NBA)
$106,900–$147,000 year
RemoteUnited States
Job Summary
Design and train reinforcement learning policies for Humana's Next Best Action platform, instrument training pipelines, and evaluate decision-making algorithms. Collaborate with data and platform engineers to ensure system operation within clinical eligibility rules and program-specific objectives. Manage software development across front-end, back-end, database integrations, and server management while influencing department strategy. Diagnose failure modes in learned policies, including instability and distributional shift, and optimize constrained sequential decisioning systems. Operate remotely with occasional travel to Tech Hubs, adhering to strict HIPAA-compliant workspace requirements and high-speed internet standards.
Required Qualifications
- 5+ years (post undergraduate level) of software engineering or quantitative research experience building and operating large-scale production systems, with emphasis on data-intensive platforms, recommendation systems, optimization engines, or simulation frameworks serving millions of users
- 2+ years (post graduate level) of software engineering or quantitative research experience building and operating large-scale production systems, with emphasis on data-intensive platforms, recommendation systems, optimization engines, or simulation frameworks serving millions of users
- 2+ years of hands-on experience implementing reinforcement learning, operations research methods, or simulation-driven decision systems in production
- Relevant backgrounds include policy gradient and value-based RL (PPO, A3C, DQN, CQL), stochastic dynamic programming, discrete-event simulation, or large-scale combinatorial or constrained optimization
- Deep familiarity with Markov Decision Processes, Bellman-equation-based value estimation, reward or objective shaping, exploration-exploitation tradeoffs, and constraint formulation in real-world decision systems
- Demonstrated ability to diagnose failure modes in learned or optimized policies: instability, poor credit assignment across long horizons, and distributional shift across large populations
- Proficiency in Python 3.x
- experience with PyTorch or TensorFlow for policy network or learned model implementation
- Experience with Ray RLlib or equivalent distributed computation frameworks for large-scale training or optimization
- Experience with Databricks, PySpark, and Delta Lake for large-scale ML or data pipelines processing tens of millions of records
- Experience with MLflow for experiment tracking, model registry, and artifact management
- Experience with shipping systems that operate reliably under production load, not just research or prototype work
- Must have the ability to provide a high-speed DSL or cable modem for a home office
- A minimum standard speed for optimal performance of 25x10 (25mpbs download x 10mpbs upload) is required
- Satellite and Wireless Internet service is NOT allowed for this role
- A dedicated space lacking ongoing interruptions to protect member PHI / HIPAA information
Desired Qualifications
- Experience with multi-agent RL frameworks (PettingZoo or equivalent) or multi-agent simulation and coordination methods
- Familiarity with operations research methods applicable to constrained sequential decisioning: linear programming, mixed-integer programming, Lagrangian relaxation, or constraint programming
- Experience operating decision or optimization systems in regulated domains (healthcare, finance, or insurance) where member safety, auditability, and explainability are requirements
- Experience building simulation environments using Gymnasium, SimPy, AnyLogic, or equivalent frameworks for policy evaluation and backtesting
- Familiarity with event-driven feedback loops and how disposition signals feed retraining or re-optimization pipelines
- OpenTelemetry instrumentation experience for ML or optimization pipeline observability
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.