8 hours ago
Base Salary
$79k - $210k/yr
Responsibilities
- Design, build, and improve highly available, self-healing systems and services with fault detection, redundancy, automated recovery, and safe handling of partial failures.
- Automate infrastructure, configuration, onboarding, and releases using infrastructure as code and continuous integration and delivery, including validation checks, staged rollouts, and recovery procedures.
- Design and execute load tests; analyze throughput, latency, resource usage, bottlenecks, and capacity improvements.
- Use telemetry, dashboards, and alerts to investigate system behavior, diagnose production issues, and convert recurring failures into tested reliability improvements.
- Develop and integrate services through secure APIs, review code, document decisions and runbooks, and collaborate with application, site reliability, security, and infrastructure teams.
Requirements
- Experience developing production software in Python, Java, or another general-purpose language, with automated testing, Git version control, and code review.
- Experience designing highly available and self-healing systems or services, including failure detection, retries, failover, and automated recovery.
- Experience with load testing and performance analysis, including latency, throughput, resource utilization, and bottleneck analysis.
- Linux troubleshooting skills covering processes, memory, storage, permissions, and networking.
- Experience developing or integrating REST APIs and working with TLS, certificates, and secure network connections.
- Experience deploying or operating cloud or containerized services using infrastructure or configuration automation.
- Clear communication and cross-team collaboration skills, including explaining tradeoffs and carrying engineering work through implementation and production validation.
- Preferred experience with telemetry, search, or streaming platforms such as OpenSearch, Elasticsearch, Logstash, or Kafka.
- Preferred experience with Kubernetes, Helm, GitOps, Oracle Cloud Infrastructure or another major cloud platform, Prometheus, Grafana, instrumentation, actionable alert design, Terraform, Salt, or Ansible.
Benefits
- Medical, dental, and vision insurance, including expert medical opinion.
- Short-term and long-term disability, life insurance, AD&D, and supplemental life insurance.
- Health care and dependent care Flexible Spending Accounts.
- Pre-tax commuter and parking benefits.
- 401(k) Savings and Investment Plan with company match.
- Paid vacation, 11 paid holidays, and paid sick leave.
- Paid parental leave and adoption assistance.
- Employee Stock Purchase Plan, financial planning, group legal, and voluntary auto, homeowner, and pet insurance.
- The role is based on the stated hiring locations and generally accepts applications for at least three calendar days or while posted.
- US-based employees must complete identity verification involving collection and processing of biometric information, subject to legally permitted accommodations.
Tech Stack
AnsibleApache KafkaElasticsearchGitGrafanaHelmJavaKubernetesLinuxLogstashOracle CloudPrometheusPythonTerraform
About Oracle
Oracle builds database technology, enterprise applications, and Oracle Cloud Infrastructure for companies and governments. Its business spans cloud subscriptions, software licenses, and support services across ERP, HCM, CX, and industry suites, plus NetSuite. Founded in 1977 and headquartered in Austin, Texas, Oracle is a public company traded on the NYSE and serves global customers migrating and running critical workloads in its cloud.
