7 days ago
Base Salary
$115k - $235k/yr
Responsibilities
- Lead development and begin architecting scalable, elastic distributed-system components.
- Define scalability requirements and optimize code and data paths for high-throughput, hyperscale workloads.
- Design fault-tolerant, durable, in-service-upgradable systems using redundancy, replication, failover, load shedding, throttling, and rate limiting.
- Define KPIs and telemetry and build dashboards, alerting, and monitoring mechanisms for system health.
- Design correctness, availability, replication, synchronization, fault-injection, brownout, performance, and load-testing scenarios.
- Diagnose and resolve production issues, participate in operational support rotations, and guide incident response and root-cause investigations.
- Implement security controls, execute remediation plans, maintain compliance documentation, and protect data and applications in multi-tenant environments.
- Develop infrastructure-as-code automation and change-management plans for patching, updating, and rolling back applications.
- Coordinate moderately complex work, provide technical oversight, mentor junior team members, and participate in candidate interviews and hiring recommendations.
Requirements
- Experience leading development and beginning architecture of scalable distributed systems.
- Ability to design high-throughput, elastic, fault-tolerant systems with replication, failover, service-level objectives, and in-service upgrades.
- Experience with telemetry, dashboards, alerting, performance and load testing, fault injection, brownouts, and operational troubleshooting.
- Knowledge of data replication and synchronization, network unreliability, consistency, availability, and partition tolerance.
- Experience implementing security controls, encryption, access controls, remediation, cloud compliance, and compliance documentation.
- Experience developing infrastructure-as-code automation and managing patching, updates, rollbacks, and change plans.
- Ability to coordinate moderately complex initiatives, provide technical oversight, collaborate across teams, solve complex problems, mentor others, and contribute to hiring.
Benefits
- Medical, dental, and vision insurance, including expert medical opinion.
- Short-term and long-term disability insurance, life insurance, AD&D, and supplemental life insurance.
- Health care and dependent care flexible spending accounts.
- Pre-tax commuter and parking benefits.
- 401(k) savings and investment plan with company match.
- Paid vacation, 11 paid holidays, and paid sick leave of 72 hours upon hire with annual refresh and carryover limits.
- Paid parental leave and adoption assistance.
- Employee Stock Purchase Plan, financial planning, group legal, and voluntary auto, homeowner, and pet insurance benefits.
- The role generally accepts applications for at least three calendar days from the posting date or while the job remains posted.
Categories
DevOpsSite Reliability
About Oracle
Oracle builds database technology, enterprise applications, and Oracle Cloud Infrastructure for companies and governments. Its business spans cloud subscriptions, software licenses, and support services across ERP, HCM, CX, and industry suites, plus NetSuite. Founded in 1977 and headquartered in Austin, Texas, Oracle is a public company traded on the NYSE and serves global customers migrating and running critical workloads in its cloud.
