
Senior Infrastructure Engineer
Venti Technologies3 days ago
Singapore, SingaporeSenior
Responsibilities
- Provide and operate a production-ready hybrid infrastructure platform for online, offline, data, streaming, platform, and internal tooling workloads.
- Design, deploy, operate, and scale Kubernetes clusters across on-premises, cloud, and hybrid environments, including capacity planning, autoscaling, upgrades, and multi-tenancy.
- Ensure high availability and strict SLO/SLA performance for central and customer-site deployments, including incident management and post-incident reviews.
- Design secure networking across on-premises, cloud, hybrid, and customer environments using VPNs, switches, routers, firewalls, and potential 5G connectivity.
- Define observability, logging, monitoring, alerting, dashboards, and distributed tracing to reduce MTTR.
- Build deployment and release automation with blue/green and canary strategies, Helm/Kustomize, GitOps, and service mesh capabilities.
- Develop internal platform capabilities, reusable infrastructure-as-code modules, golden paths, and self-service provisioning.
- Establish operational processes for change management, capacity planning, security patching, disaster recovery, and runbooks.
Requirements
- Bachelor’s or Master’s degree in Computer Science or a related field.
- At least 5 years of experience implementing and operating on-premises and cloud compute, storage, and networking infrastructure.
- At least 3 years of experience designing and operating production Kubernetes or OpenShift clusters at scale, including autoscaling, networking, storage, security, and upgrades.
- Strong Linux administration and scripting skills with hands-on Bash, Python, and Terraform experience.
- Experience with hybrid infrastructure across on-premises, Azure, AWS, or GCP.
- Experience with Terraform, Ansible, Packer, ArgoCD or Flux, Helm, and Kustomize for infrastructure automation and configuration management.
- Production experience running Docker and Kubernetes or OpenShift at scale.
- Experience designing and operating observability stacks using Prometheus/Grafana, ELK/Loki/OpenSearch, and OpenTelemetry.
- Strong understanding of infrastructure security, including network segmentation, mTLS, RBAC, secrets management, image signing, vulnerability management, and CIS benchmarks.
- Hands-on experience with on-premises VPNs, switches, routers, firewalls, and network troubleshooting.
- Excellent communication and collaboration skills, including work with cross-functional teams and customer-facing deployment environments.
- Bonus experience with OpenStack, internal developer platforms, service mesh at multi-cluster scale, high-throughput video streaming, 5G connectivity, autonomous driving, robotics, or edge environments.
Benefits
- World-class benefits and a collaborative international working environment.
- Flexible working arrangements.
Tech Stack
AnsibleAWSAzureBashDockerGoogle Cloud PlatformGrafanaHelmKubernetesLinuxOpenShiftOpenStackPrometheusPythonTerraform
Categories
DevOpsSite Reliability