One Identity logo
One IdentityPosted 1 month ago

Senior Software Engineer – Production Engineering

RemoteBudapest, Budapest, Hungary or Hungary

Full TimeSenior LevelMedium

Job Summary

Debug complex production issues, strengthen observability across logs, metrics, and alerting, and eliminate recurring problems through code fixes and architectural improvements. Improve mean time to detect and resolve incidents while designing debugging workflows, runbooks, and internal tooling to reduce operational burden. Partner with product teams to feed production learnings back into design, ensuring systems are easier to understand, operate, and troubleshoot. This hands-on role requires a 24/7 on-call rotation and utilizes AI-augmented development to accelerate root cause identification. You will operate as a T-shaped engineer deep in production reliability while maintaining capability across backend, APIs, and cloud environments.

Required Qualifications

  • 4+ years of software engineering experience with ownership of production systems, reliability, or operational improvements
  • Strong backend development experience (Ruby, Node.js, or similar)
  • Solid understanding of REST APIs, service contracts, and software design principles
  • Experience working across backend services, APIs, and production systems
  • Experience building and operating services in AWS or similar cloud environments
  • Good understanding of distributed systems, cloud-native architecture, and CI/CD
  • Experience with observability, production debugging, and incident response
  • Willingness to participate in a mandatory 24/7 on-call rotation
  • Experience responding to production incidents and contributing to reliability improvements
  • Experience using, or strong interest in, AI-powered development tools (e.g., GitHub Copilot, ChatGPT, Cursor)
  • Ability to evaluate and refine AI-generated code

Desired Qualifications

  • Strong production debugging experience in distributed systems
  • Experience troubleshooting complex, customer-impacting issues under real-world conditions
  • Deep familiarity with observability tools (Datadog or similar — logs, metrics, APM, tracing)
  • Experience improving operational workflows (runbooks, incident response, debugging tooling)
  • Ability to identify patterns across incidents and drive systemic fixes, not just one-off resolutions
  • Experience working across service boundaries (APIs, databases, infrastructure) to diagnose issues

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce