Operations Engineer II
RemoteNew Jersey, United States or North Carolina, United States
New Jersey, United States or North Carolina, United StatesRemoteFull TimeBachelors DegreeMarketing SoftwareSmall
Full TimeBachelors DegreeSmallMarketing Software
Job Summary
Design and build solutions for AI agents to implement infrastructure, reviewing generated code for quality and correctness. Provision production systems, tune monitoring and alerting, and coordinate work across teams to maintain platform reliability. Define infrastructure as code in Terraform, validate Kubernetes changes, wire up Grafana alerts, and manage releases through GitHub and deployment pipelines. Participate in the 24x7 on-call rotation while generating documentation for both humans and AI agents.
Required Qualifications
- Bachelor's degree in a STEM-related field or equivalent work experience
- 2+ years in a platform engineering or systems/operations role
- Working knowledge of AI development tools and coding assistants, with comfort using them in day-to-day implementation work
- Strong analytical and troubleshooting skills — you fix root causes, not just symptoms — paired with the communication skills to translate technical issues for different audiences
- Experience in a 24x7 on-call rotation (incident response, mitigation, root cause analysis)
- Troubleshooting across Linux and Mac-based environments
- Using Git or Mercurial for version control
- Experience with containerized, cloud-based infrastructure — Kubernetes, Docker, and infrastructure-as-code tools like Terraform or CloudFormation on AWS, GCP, or similar
- Comfort with shell scripting and Python
- Experience configuring monitoring and alerting with tools like Grafana, Prometheus, PagerDuty, or Sentry
- Solid grasp of networking fundamentals (DNS, VPN, load balancing, subnetting) and security concepts (firewalls, secrets management, IAM, PKI), with practical experience applying them
Desired Qualifications
- Bachelor's degree in a STEM-related field or equivalent work experience, plus 2+ years in a platform engineering or systems/operations role
- Working knowledge of AI development tools and coding assistants, with comfort using them in day-to-day implementation work
- Strong analytical and troubleshooting skills — you fix root causes, not just symptoms — paired with the communication skills to translate technical issues for different audiences
- Experience in a 24x7 on-call rotation (incident response, mitigation, root cause analysis), troubleshooting across Linux and Mac-based environments, and using Git or Mercurial for version control
- Experience with containerized, cloud-based infrastructure — Kubernetes, Docker, and infrastructure-as-code tools like Terraform or CloudFormation on AWS, GCP, or similar
- Comfort with shell scripting and Python, and experience configuring monitoring and alerting with tools like Grafana, Prometheus, PagerDuty, or Sentry
- Solid grasp of networking fundamentals (DNS, VPN, load balancing, subnetting) and security concepts (firewalls, secrets management, IAM, PKI), with practical experience applying them
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.