Oracle logo
OraclePosted 1 month ago

Principal Core Infrastructure Engineer (OCI Object Storage)

On-siteSanta Clara, California, United States

Full TimeSenior LevelSmallCloud Services

Job Summary

Lead development and architect scalable, elastic distributed systems for the Object Storage Service, defining scalability requirements and optimizing code paths for high-throughput, hyper-scale workloads. Design fault-tolerant, in-service-upgradable systems using redundancy, replication, and failover policies while applying load-shedding and rate-limiting to meet SLOs. Establish KPIs and telemetry, build proactive dashboards, and design complex validation including fault injection and synchronization for correctness and durability. Proactively diagnose and resolve production issues, mentor peers, and implement robust security controls with IaC automation to enable safe patching and rollbacks within change-management plans. Drive the next phase of best-in-class Object Storage features for enterprise and big data workloads.

Required Qualifications

  • expertise and passion in solving difficult problems in distributed systems
  • large scale storage
  • highly available services
  • familiarity of distributed systems
  • value simplicity and scale
  • work comfortably in a collaborative, agile environment
  • be excited to learn
  • ability to lead development and begin architecting components of scalable, elastic distributed systems
  • ability to define and enforce scalability requirements for owned components
  • ability to optimize code and data paths for high-throughput, hyper-scale workloads
  • ability to leverage data plane platforms for large-scale retrieval, storage, and processing
  • ability to design fault-tolerant, in-service-upgradable systems using redundancy, replication, failover, and policies for partitions
  • ability to apply load-shedding, throttling, and rate-limiting to handle network unreliability while meeting SLOs
  • ability to establish KPIs and telemetry
  • ability to build proactive dashboards and alerts
  • ability to design complex validation (fault injection, brownouts), replication, and synchronization for correctness and durability
  • ability to proactively diagnose and resolve production issues
  • ability to mentor peers
  • ability to ensure operational readiness
  • ability to implement robust security controls
  • ability to execute remediation
  • ability to maintain compliance documentation
  • ability to develop IaC and automation that enable safe patching, updates, and rollbacks within change-management plans

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce