JPMorgan Chase & Co logo
JPMorgan Chase & CoPosted 1 month ago

Principal Software Engineer - Developer Platform Engineering

On-sitePalo Alto, California, United States

Full TimeSenior LevelEnterpriseFinancial Services

Job Summary

Define and own a portfolio-wide capacity evaluation methodology through workload modeling, bottleneck identification, and automated assessment at scale. Build a developer platform engineering practice by delivering paved roads, self-service tooling, and templated performance test harnesses for adoption. Set technical direction for performance, capacity, and resiliency standards, including SLOs, error budgets, and performance budgets, driving adoption through scorecards and release-readiness gates. Design and implement automated performance evaluation pipelines embedded in CI/CD with regression detection and actionable reporting. Lead resiliency and chaos engineering strategy to validate autoscaling behavior and graceful degradation patterns under extreme load. Coach and enable performance engineers and application teams on performance-first design, shifting testing and validation earlier through self-service frameworks. Architect and govern agentic AI-enabled engineering workflows to improve delivery speed and operational outcomes while defining guardrails for validation, security, and resiliency.

Required Qualifications

  • 15+ years of overall engineering experience
  • 8+ years in performance engineering, capacity planning, or developer platform engineering for high-traffic distributed systems
  • Demonstrated experience scaling engineering practices across large organizations (500+ engineers) through self-service platforms, maturity models, and adoption playbooks
  • Formal training or certification on software engineering concepts
  • 10+ years applied experience
  • Hands-on software engineering proficiency in Java and Spring Boot
  • Strong familiarity in modern front-end ecosystems such as React and server-side rendering
  • Deep expertise in workload modeling and capacity estimation, including statistical analysis of latency and throughput and translating portfolio goals into measurable service-level outcomes
  • Strong application performance monitoring and observability experience (for example, Dynatrace or OpenTelemetry), including distributed tracing across request chains and continuous profiling approaches
  • Experience embedding performance automation into CI/CD pipelines with regression detection, environment-aware test execution, and enforceable quality gates
  • Proficiency with load testing tools (for example, JMeter, k6, or Gatling) and service virtualization or fault injection approaches for dependency isolation
  • Kubernetes and cloud platform experience (for example, Amazon Elastic Kubernetes Service on Amazon Web Services), including autoscaling strategies, infrastructure as code, and cost-per-transaction awareness
  • Demonstrated experience designing and leading adoption of agentic AI-enabled development practices (using enterprise-authorized tools within the work environment) across teams, including setting standards for human-in-the-loop validation, auditability/traceability of changes, and secure handling of sensitive data
  • Strong understanding of responsible AI use and control expectations in engineering workflows, including security/resiliency implications, data sensitivity, and risk-based governance; ability to influence senior technical leaders on safe scaling patterns and reuse

Desired Qualifications

  • Experience with front-end performance optimization, including Core Web Vitals, Lighthouse CI, bundle analysis, and browser-based synthetic testing (for example, Playwright or WebPageTest)
  • Experience with data platform and payment system performance, such as database tuning, event streaming throughput optimization, batch processing profiling, and rate-limit validation
  • Strong systems and cloud performance background, including Linux tooling, Java Virtual Machine tuning, containers, networking, and modern HTTP optimization
  • Familiarity with service mesh technologies and traffic-control patterns such as backpressure, circuit breaking, and cost-efficiency analysis
  • Practical application of large language models and agent-based automation for workload generation, anomaly detection, capacity report automation, or test maintenance at scale

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce