17 days ago
Base Salary
$115k - $235k/yr
Responsibilities
- Lead the design, implementation, and evolution of hyperscale distributed systems and data-plane services.
- Define and enforce scalability, elasticity, durability, availability, and reliability requirements.
- Optimize high-throughput data paths using distributed state, replication, and synchronization patterns.
- Design fault-tolerant systems with redundancy, automatic failover, recovery, load shedding, throttling, retries, and timeouts.
- Establish SLOs, KPIs, telemetry, dashboards, and proactive alerting for critical systems.
- Lead performance, load, fault-injection, and brownout testing.
- Diagnose production incidents, guide root-cause analysis, and mentor engineers in operational excellence.
- Build Infrastructure as Code and operational automation for patching, updates, rollbacks, and change management.
- Apply encryption, access controls, security remediation, and compliance practices for multi-tenant cloud infrastructure.
- Automate GPU test validation, configuration management, repair and triage, manufacturing quality metrics, and deployment-readiness reporting.
- Collaborate with supply chain, hardware development, manufacturing partners, data-center operations, NVIDIA, and AMD.
Requirements
- Bachelor's or master's degree in Computer Science, Computer Engineering, or a related field, or equivalent practical experience.
- At least 7 years of professional software-engineering experience with large-scale distributed systems or cloud infrastructure.
- Strong experience designing and operating highly available, scalable, and fault-tolerant distributed systems.
- Proficiency in one or more of Java, C++, C#, or Go.
- Deep knowledge of distributed-systems design, data structures, algorithms, operating systems, networking, and secure software development.
- Experience with system-level test automation, performance and load testing, reliability engineering, and production incident response.
- Experience leading or influencing technical architecture and mentoring engineers.
- Preferred experience with Oracle Cloud, AWS, Azure, Google Cloud, distributed storage, microservices, replication, state management, observability, Infrastructure as Code, security controls, compliance, and 24x7 production operations.
Benefits
- Medical, dental, vision, disability, life, and AD&D insurance.
- Flexible spending accounts, pre-tax commuter and parking benefits, and a 401(k) plan with company match.
- Paid vacation, 11 paid holidays, paid sick leave, paid parental leave, and adoption assistance.
- Employee Stock Purchase Plan, financial planning, group legal benefits, and voluntary insurance benefits.
- U.S. base salary range of $114,600 to $234,600 per annum, with possible bonus, equity, and compensation deferral.
- U.S.-based employees complete identity verification involving biometric information, subject to applicable accommodations.
About Oracle
Oracle builds database technology, enterprise applications, and Oracle Cloud Infrastructure for companies and governments. Its business spans cloud subscriptions, software licenses, and support services across ERP, HCM, CX, and industry suites, plus NetSuite. Founded in 1977 and headquartered in Austin, Texas, Oracle is a public company traded on the NYSE and serves global customers migrating and running critical workloads in its cloud.
