8 days ago
Base Salary
$115k - $235k/yr
Responsibilities
- Lead the design, implementation, and evolution of core distributed systems and data-plane services at hyperscale.
- Define scalability, elasticity, durability, availability, and reliability requirements for owned components.
- Optimize high-throughput data paths for retrieval, storage, and processing using distributed state, replication, and synchronization.
- Design fault-tolerant systems with redundancy, automatic failover, in-service updates, and recovery-oriented behavior.
- Implement reliability patterns including load shedding, throttling, rate limiting, retries, and timeouts.
- Establish SLOs, KPIs, telemetry, dashboards, and proactive alerting for critical systems.
- Lead performance, load, fault-injection, and brownout testing for correctness, resilience, and operational readiness.
- Diagnose and recover from production incidents, lead root-cause analysis, and mentor engineers in operational excellence.
- Build Infrastructure as Code and operational automation for patching, updates, rollbacks, and change management.
- Apply encryption, access controls, remediation practices, and compliance-ready security controls to multi-tenant cloud infrastructure.
- Automate GPU test validation and maintain consistent firmware, software, and hardware configurations.
- Develop repair and triage capabilities and establish manufacturing yield, throughput, quality, and deployment-readiness metrics.
Requirements
- Bachelor's or master's degree in Computer Science, Computer Engineering, or a related field, or equivalent practical experience.
- At least 7 years of professional software-engineering experience with demonstrated impact on large-scale distributed systems or cloud infrastructure.
- Strong experience designing and operating highly available, scalable, and fault-tolerant distributed systems.
- Proficiency in one or more object-oriented or systems programming languages such as Java, C++, C#, or Go.
- Deep understanding of distributed-systems design, data structures, algorithms, operating systems, networking, and secure software-development practices.
- Experience with system-level test automation, performance and load testing, reliability engineering, and production incident response.
- Demonstrated experience leading or influencing technical architecture and mentoring engineers.
- Preferred experience with Oracle Cloud, AWS, Azure, Google Cloud, data-plane platforms, distributed storage, microservices, replication, state management, or high-throughput data processing.
- Preferred experience defining SLOs, building observability systems, operating 24x7 production services, and implementing infrastructure automation, security controls, and cloud compliance requirements.
- Strong problem-solving, communication, and cross-functional collaboration skills.
Benefits
- US hiring range is $114,600 to $234,600 per annum, with possible bonus, equity, and compensation deferral.
- Benefits include medical, dental, vision, disability, life insurance, flexible spending accounts, commuter and parking benefits, 401(k) matching, paid time off, paid holidays, paid sick leave, parental leave, adoption assistance, employee stock purchase, financial planning, group legal, and voluntary insurance benefits.
- US-based employees must complete identity verification involving collection and processing of biometric information, subject to legally required accommodations.
About Oracle
Oracle builds database technology, enterprise applications, and Oracle Cloud Infrastructure for companies and governments. Its business spans cloud subscriptions, software licenses, and support services across ERP, HCM, CX, and industry suites, plus NetSuite. Founded in 1977 and headquartered in Austin, Texas, Oracle is a public company traded on the NYSE and serves global customers migrating and running critical workloads in its cloud.
