Wealthsimple logo
WealthsimplePosted 1 month ago

Staff Software Engineer, Observability Platform

Remote

Full TimeSenior LevelMedium

Job Summary

Design and build the events-first foundation, owning the architecture of wide-event data models and high-throughput ingest, storage, and query pipelines. Develop instrumentation libraries, shared SDKs, and golden paths to accelerate telemetry adoption across engineering teams. Define technical standards for OpenTelemetry, naming, tagging, and context propagation to ensure humans and AI agents share a unified language. Make production legible by constructing fast query foundations and protocols like MCP that enable AI agents to investigate incidents and operate alongside engineers. Use AI coding tools fluently to prototype, build, and navigate large systems while raising the team's bar for AI-assisted development. Mentor senior engineers, review complex designs, and lead cross-team initiatives through influence.

Required Qualifications

  • Significant software engineering experience, typically 8+ years
  • Software development background rather than a primarily operations or systems-administration one
  • Hands-on, current coder in more than one language such as Kotlin and Ruby
  • Comfort across multiple stacks
  • Meaningful time writing and shipping production code
  • Direct experience building an observability or telemetry platform and its SDKs
  • Shipped instrumentation libraries, pipelines, and shared platforms
  • Deep technical expertise in observability
  • Depth in event-based and high-cardinality systems
  • Understanding of wide structured events
  • Understanding of columnar or streaming backends
  • Understanding of cardinality
  • Understanding of sampling
  • Understanding of tagging
  • Understanding of context propagation
  • Understanding of tradeoffs of ingesting and querying telemetry at scale
  • Fluency with open standards and instrumentation
  • Hands-on experience with OpenTelemetry or an equivalent
  • Experience with SLI and SLO design
  • Comfort making significant changes in ambiguous codebases
  • Ability to enter a system you did not build
  • Ability to form an accurate model of a system quickly
  • Ability to change a system safely and substantially
  • Proficiency with AI coding tools and LLMs
  • Fluency with AI coding tools such as Claude Code
  • Clear point of view on where AI tools help and where they do not
  • Strong communication and influence
  • Ability to partner well with other engineering teams
  • Ability to explain complex tradeoffs clearly
  • Ability to improve the systems and the people around you
  • Ability to work confidently in ambiguity
  • Ability to jump into unfamiliar codebases
  • Ability to make significant, well-reasoned changes with high impact
  • Ability to prove value through experiments
  • Ability to pilot new approaches with one or two teams
  • Ability to measure results of pilots
  • Ability to scale what works rather than committing everything up front
  • Ability to raise the bar for others
  • Ability to mentor senior engineers
  • Ability to review complex designs
  • Ability to lead cross-team technical initiatives through influence rather than authority
  • Ability to design the foundations
  • Ability to write the code that matters most
  • Ability to raise the engineering quality of everyone around you
  • Ability to design and build the events-first foundation
  • Ability to own the architecture of the wide-event data model
  • Ability to own the high-throughput ingest, storage, and query pipelines
  • Ability to handle high-cardinality and columnar or streaming systems
  • Ability to design and ship instrumentation libraries
  • Ability to design and ship shared SDKs
  • Ability to drive adoption of rich, consistent telemetry
  • Ability to define instrumentation conventions
  • Ability to define naming practices
  • Ability to define tagging practices
  • Ability to define sampling practices
  • Ability to define context-propagation practices
  • Ability to codify practices on open standards
  • Ability to build the fast query foundation
  • Ability to build access patterns including protocols such as MCP
  • Ability to let AI agents investigate incidents
  • Ability to let AI agents verify their own changes
  • Ability to let AI agents operate in tight feedback loops alongside engineers
  • Ability to use AI coding tools such as Claude Code fluently to prototype
  • Ability to use AI coding tools such as Claude Code fluently to build
  • Ability to use AI coding tools such as Claude Code fluently to navigate large systems
  • Ability to help the team raise its own bar for building with AI
  • Ability to partner well with other engineering teams
  • Ability to explain complex tradeoffs clearly
  • Ability to improve the systems and the people around you

Desired Qualifications

  • Experience in a regulated or fintech environment
  • Familiarity with Kubernetes
  • Experience with progressive delivery such as Argo Rollouts
  • Experience with columnar stores such as ClickHouse
  • Experience building observability for LLM and agent-based workloads

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce