Level

Senior Infrastructure Engineer

Level
Apply
2 hours ago
Auckland, New ZealandSenior

Responsibilities

  • Design, build, and operate secure, highly available AWS infrastructure with Terraform or OpenTofu and Atlantis-based delivery.
  • Operate and evolve the EKS platform, including autoscaling with Karpenter, upgrades, core add-ons, and Helm-based delivery through ArgoCD.
  • Build and maintain GitHub Actions CI/CD pipelines and self-service developer tooling.
  • Own ingress and egress, service mesh, mTLS, load balancing, edge TLS, and DNS using technologies including Traefik, Linkerd, Envoy, and Route 53.
  • Build observability with OpenTelemetry and SigNoz and use telemetry to guide reliability, performance, and cost decisions.
  • Serve as an escalation point for complex incidents and lead troubleshooting and post-mortems.
  • Apply cloud security practices across identity, secrets, network boundaries, vulnerability tooling, and organizational guardrails.
  • Lead infrastructure projects independently, set standards, mentor engineers, and improve the platform continuously.

Requirements

  • 5+ years of experience operating large-scale cloud infrastructure, with AWS strongly preferred.
  • Deep infrastructure-as-code experience with Terraform or OpenTofu; CloudFormation, Pulumi, or CDK experience is also relevant.
  • Strong production experience with Docker and Kubernetes, including EKS.
  • Scripting and automation proficiency in Python, Go, or Bash.
  • Solid cloud networking knowledge covering VPC, DNS, load balancing, ingress, firewalls/WAF, and VPNs.
  • Proven CI/CD and GitOps experience with GitHub Actions or a similar tool.
  • Experience using metrics, logs, and traces to drive operational decisions.
  • Demonstrated ability to lead infrastructure projects independently from end to end.
  • Strong communication skills across technical and non-technical audiences.
  • Preferred experience with ArgoCD, Atlantis, Linkerd or Envoy, Traefik, Karpenter, Helm, OpenTelemetry, SigNoz, Grafana, Datadog, Prometheus, Backstage, or other internal developer platforms.
  • AI/ML infrastructure experience involving GPU scheduling, model or agent hosting, or inference gateways is preferred.
  • Rust service CI/CD, CloudFront/CDN, distributed systems, or open-source infrastructure contributions are preferred.
  • AWS Solutions Architect or DevOps Engineer Professional certification is preferred.

Tech Stack

AmbassadorAWSBashDatadogDockerGitHub ActionsGoGrafanaHelmKubernetesPrometheusPythonRustTerraform

Categories

Level

About Level

201-500 employees

Level builds a learning technology platform that combines established curriculum principles with interactive, game-style design to help students practice academic and life skills. The product runs on a forked, optimized Godot engine engineered to perform reliably on mobile and other constrained hardware. It targets students and educators seeking repeatable, meaningful practice in a format that feels like a game.

Contact me