NinjaTrader logo
NinjaTraderPosted 1 week ago

Staff Site Reliability Engineer

$160,000–$210,000 year

On-siteChicago, Illinois, United States

Full TimeSenior LevelMediumFinance Software

Job Summary

Deploy single- and multi-cluster services to Kubernetes, partnering with Product Engineering and Technical Operations to strengthen monitoring, reduce operational toil, and maintain 99.95% uptime across revenue-generating systems. Serve as the technical lead for the SRE function, setting direction and mentoring engineers while performing initial root cause analysis and remediation of production incidents. Build automation tools for deployments and incident responses, design reliable monitoring systems with SLIs/SLOs, and leverage Infrastructure as Code to manage cloud environments. Participate in a weekly 12x7 on-call rotation, including weekend deployments and Sunday market checkouts. Join the Platform Engineering team to advocate for customers and ensure the availability of NinjaTrader's trading platform.

Required Qualifications

  • 8+ years of experience in DevOps, Site Reliability Engineering, or Platform Engineering roles
  • Expertise with Kubernetes, Docker, and container orchestration
  • Hands-on experience with CI/CD tools (GitHub Actions or equivalent)
  • Proficiency in programming languages (e.g., Python, Bash, or Go) and automation tools such as Ansible, Terraform, or Helm
  • Hands-on experience with AWS, GCP, or Azure, including in-depth knowledge of networking, security, and identity management in cloud environments
  • Knowledge of monitoring and observability tools such as Prometheus, Grafana, Datadog, or similar
  • Strong collaboration, communication, and leadership skills, with the ability to influence technical decisions across teams and mentor junior engineers
  • Participate in a weekly 12x7 on-call rotation, including weekend deployments and checkouts before markets open on Sundays
  • Serve as the technical lead for the SRE function, setting technical direction and mentoring engineers across reliability initiatives
  • Analyze, troubleshoot, and remediate production issues with a systematic problem-solving approach to keep revenue-generating systems running
  • Perform initial root cause analysis and remediation of production incidents
  • Build tools that automate repetitive tasks, deployments, and incident responses with the goal of minimal human involvement
  • Design reliable monitoring and alerting systems with Product and QA teams, establishing and tracking SLIs/SLOs across web, mobile, desktop, and trading platforms
  • Leverage Infrastructure as Code tools like Terraform to automate provisioning, scaling, and management across all platforms
  • Collaborate with Product Engineering, Operations, and other cross-functional teams to deliver features on time while meeting scalability, security, and performance requirements
  • Implement security and compliance best practices (e.g., SOC 2, PCI DSS) throughout the software delivery lifecycle for web, mobile, and desktop deployments
  • Must be based in Chicago, IL
  • Must be available for hybrid work schedule: In-office Tuesday through Thursday, with remote work on Mondays and Fridays

Desired Qualifications

  • Trading industry experience
  • Contributions to open-source projects

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce