11 days ago
Bengaluru, IndiaStaff+
Responsibilities
- Lead architecture and design of elastic, hyper-scale distributed systems and data plane platforms.
- Identify and resolve performance, scalability, reliability, and correctness challenges in high-throughput systems.
- Design fault tolerance, redundancy, replication, automatic failover, service-disruption handling, and in-service upgrades.
- Define KPIs, telemetry, dashboards, alerting, SLO-aligned availability, and durability standards.
- Apply formal verification techniques such as TLA+ and develop data replication and synchronization strategies.
- Advise on complex production troubleshooting, lead incident response and root-cause investigations, and establish operational readiness and SOP standards.
- Architect security controls for multi-tenant environments, guide vulnerability remediation, and ensure cloud infrastructure compliance.
- Develop enterprise automation and Infrastructure as Code strategies for safe patching, updates, rollbacks, and change management.
- Manage or provide direction on timelines, deliverables, budgets, cross-functional alignment, and strategic initiatives.
- Mentor engineers, contribute to talent development and hiring decisions, and provide technical thought leadership.
Benefits
- Competitive benefits including flexible medical, life insurance, and retirement options.
- Volunteer programs supporting employee community involvement.
- Accessibility assistance and accommodation support during the employment process.
Categories
DevOpsSite Reliability
About Oracle
Oracle builds database technology, enterprise applications, and Oracle Cloud Infrastructure for companies and governments. Its business spans cloud subscriptions, software licenses, and support services across ERP, HCM, CX, and industry suites, plus NetSuite. Founded in 1977 and headquartered in Austin, Texas, Oracle is a public company traded on the NYSE and serves global customers migrating and running critical workloads in its cloud.
