Oracle

Principal Core Infrastructure Engineer – GPU

Oracle
Apply
10 hours ago
Hong Kong, Hong KongStaff+

Responsibilities

  • Design, deploy, and manage cloud resources, distributed computing systems, data storage, and GPU infrastructure supporting AI/ML workflows.
  • Collaborate with customer data scientists, software engineers, infrastructure teams, and internal cloud services teams on training, testing, and production deployments.
  • Implement automation for provisioning, configuration, monitoring, and operational productivity.
  • Optimize infrastructure performance, resource utilization, caching, data preprocessing, reliability, scalability, and cost-effectiveness.
  • Troubleshoot proof-of-concept and production infrastructure issues and implement solutions to reduce risk and downtime.
  • Ensure encryption, access control, vulnerability management, security, and compliance standards across the AI/ML infrastructure stack.
  • Advise customers and partners on advanced GPU solutions, cloud migration, Oracle Cloud architecture, and deployment best practices.
  • Lead technical and architectural discussions and serve as a liaison between customers, core engineering teams, and support.
  • Document infrastructure designs, configurations, procedures, and troubleshooting guidance.
  • Coordinate moderately complex projects and initiatives, monitor deliverables, prioritize work, and provide technical oversight.
  • Coach and mentor junior team members, contribute to process improvements, and participate in candidate interviews and hiring recommendations.

Requirements

  • Experience with scripting and automation using Ansible, Terraform, Python, and/or Kubernetes.
  • Experience with Docker, Kubernetes, and distributed-system orchestration tools such as Slurm or PBS.
  • Understanding of networking concepts, security principles, and security best practices.
  • Strong troubleshooting, problem-solving, communication, collaboration, and technical documentation skills.
  • Hands-on Linux administration experience with Oracle Linux, RHEL, CentOS, Ubuntu, and Debian, including package management, shell scripting, and performance optimization.
  • Strong proficiency in at least one of Python, Rust, Go, Java, or Scala.
  • Experience designing, implementing, and managing infrastructure for AI/ML or HPC workloads.
  • Understanding of TensorFlow, PyTorch, or scikit-learn and their deployment in production environments is preferred.
  • Familiarity with DevOps tools and practices, including Jenkins, GitLab CI/CD, and Prometheus, is preferred.
  • Strong experience with high-performance computing and GPU systems.
  • Ability to think strategically about business, products, and technical challenges and work effectively with large customers.
  • Ability to coach junior team members and contribute to candidate assessment and hiring recommendations.

Benefits

  • Competitive benefits including flexible medical, life insurance, and retirement options.
  • Employee volunteer programs and opportunities to give back to the community.
  • Equal employment opportunity and disability accommodation support are provided.

Tech Stack

AnsibleDockerGitLab CI/CDGoJavaJenkinsKubernetesOracle CloudPrometheusPythonPyTorchRustScalascikit-learnTensorFlowTerraform

Categories

Forward Deployed
Oracle

About Oracle

10,000+ employees

Oracle builds database technology, enterprise applications, and Oracle Cloud Infrastructure for companies and governments. Its business spans cloud subscriptions, software licenses, and support services across ERP, HCM, CX, and industry suites, plus NetSuite. Founded in 1977 and headquartered in Austin, Texas, Oracle is a public company traded on the NYSE and serves global customers migrating and running critical workloads in its cloud.

Contact me