Base Salary
$105k - $235k/yr
Responsibilities
- Define the technical vision, architecture, and multiyear evolution of OCI’s AI-native developer platforms.
- Design and deliver secure, highly available services for large-scale distributed cloud environments.
- Architect systems for service-behavior analysis, traffic modeling, realistic workload generation, and functional and non-functional evaluation.
- Establish safety-controlled agent-assisted workflows for canary, functional, integration, load, and performance testing.
- Apply machine learning and LLM capabilities to telemetry, API changes, incidents, test results, and engineering knowledge.
- Develop approaches for failure diagnosis, regression detection, workload modeling, and converting production incidents into reusable test coverage.
- Define platform APIs, data models, extension points, and integration patterns for adoption across OCI services.
- Set standards for scalability, reliability, observability, security, privacy, performance, and responsible AI.
- Define metrics and feedback loops for test effectiveness, traffic coverage, developer effort saved, reliability, and adoption.
- Provide hands-on technical leadership, mentor senior engineers, resolve cross-team architecture issues, and improve engineering standards.
Requirements
- Bachelor’s or master’s degree in computer science, engineering, or a related technical field, or equivalent practical experience.
- At least eight years of professional software engineering experience, including significant experience designing and operating large-scale production systems.
- Proficiency in one or more modern programming languages such as Java, Go, Python, or C++.
- Deep expertise in distributed systems, system design, APIs, concurrency, data modeling, and scalable service architectures.
- Experience delivering complex platforms, backend services, microservices, event-driven systems, or large-scale data pipelines.
- Experience developing and operating cloud-native software with CI/CD, containerization, orchestration, and infrastructure automation.
- Strong understanding of service reliability, observability, performance engineering, security, and production operations.
- Experience integrating machine learning, generative AI, or other intelligent capabilities into production software.
- Ability to lead high-impact initiatives across teams, influence without direct authority, and navigate ambiguous requirements.
- Strong technical judgment and communication skills for explaining architecture to technical, product, and executive stakeholders.
- Experience mentoring senior engineers and improving engineering practices across teams or organizations.
- Preferred: experience with developer platforms, CI/CD systems, testing infrastructure, engineering productivity tools, ML or generative-AI production solutions, AI agents, retrieval-augmented generation, prompt engineering, LLM-as-a-judge methods, or model evaluation.
- Preferred: experience with PyTorch, TensorFlow, Hugging Face Transformers, traffic replay, workload modeling, canary analysis, load or performance testing, chaos engineering, Kubernetes, containers, service meshes, OCI or another major public cloud, multi-tenant platforms, API schemas, service dependencies, incidents, traffic patterns, AI safety, model monitoring, privacy, and secure enterprise data handling.
Benefits
- Comprehensive medical, dental, and vision insurance, including expert medical opinion.
- Short-term and long-term disability, life insurance, AD&D, and supplemental life insurance.
- Health care and dependent care flexible spending accounts.
- Pre-tax commuter and parking benefits.
- 401(k) savings and investment plan with company match.
- Paid vacation, 11 paid holidays, and paid sick leave.
- Paid parental leave and adoption assistance.
- Employee Stock Purchase Plan.
- Financial planning, group legal, and voluntary auto, homeowner, and pet insurance benefits.
- The role is a U.S.-based salaried position with applications generally accepted for at least three calendar days or while posted.
Categories
About Oracle
Oracle is a global leader in AI, delivering the cloud infrastructure, data, and applications that organizations across the world trust to successfully achieve business outcomes at scale. Oracle Cloud Infrastructure (OCI) provides fast, flexible, scalable AI infrastructure. With superior compute performance and network design, a comprehensive choice of AI services for developing and orchestrating agentic AI workflows at scale, and unrivaled data control, security, privacy, and governance, OCI is designed for AI workloads. It also gives customers the flexibility to run their workloads wherever they need by supporting public cloud, multicloud, hybrid cloud, sovereign, and on-premises deployments. The Oracle AI Database is the only database with AI natively built in everywhere. By bringing AI to where data lives, the Oracle AI Database delivers enterprise-grade AI securely, cost-effectively, and at scale. In-database machine learning and vector search enable AI models to run where data resides. Support for multi-modal data (structured, unstructured, graph, vector) enables richer AI use cases. The Oracle AI Data Platform enables organizations to build, deploy, and govern AI agents and applications on top of private, secure, and distributed enterprise data. This enables organizations to operationalize AI without compromising data security or control. Oracle Applications embed AI directly into business workflows and industry processes to ensure AI is delivered in context, where work happens. They include the only applications suite that brings AI to every aspect of a business, with AI-powered ERP, HCM, and CX applications, and the Oracle Health Suite which provides applications for managing the entire healthcare ecosystem. Oracle also provides industry-specific cloud application suites for more than two dozen industries and NetSuite, the world’s first cloud computing company. This LinkedIn page serves as Oracle’s official global company page.
