The Trade Desk

Sr Site Reliability Engineer

The Trade Desk
Apply
5 months ago
Sydney, AustraliaSenior
H1B Sponsor

Responsibilities

  • Design, build, and scale a global network platform across physical datacenters and AWS, Azure, and Alibaba Cloud.
  • Support thousands of hosts worldwide and engineer reliable solutions for petabyte-scale data challenges.
  • Troubleshoot and resolve complex network issues while maintaining high availability and performance.
  • Lead root cause analyses and postmortems and convert incidents into operational improvements.
  • Build tools, automate workflows, and eliminate operational toil.
  • Participate in a global follow-the-sun on-call rotation.
  • Partner with SRE and infrastructure teams to shape network automation strategy and build scalable, maintainable solutions.

Requirements

  • 6–8 years of hands-on network automation and operational experience supporting large-scale production infrastructure.
  • Strong software development and networking experience with a software-first mindset.
  • Deep expertise in TCP/IP, the OSI model, BGP, OSPF, and large-scale IP networking.
  • Hands-on experience with Kubernetes networking technologies such as Cilium and Calico and understanding of CNIs.
  • Experience managing NGINX Ingress, Envoy, or HAProxy in large-scale production environments.
  • Experience troubleshooting and performance tuning Kubernetes and Docker networking; bare-metal Kubernetes experience is a plus.
  • Knowledge of IPv6, SDN, SDN controllers, QoS, and bandwidth management.
  • Experience operating SONiC, Cisco IOS, JunOS, Arista EOS, or Nokia SR Linux/SR OS at scale.
  • Experience with monitoring, alerting, complex rules, and time-series queries using tools such as Prometheus and Grafana.
  • Experience applying infrastructure-as-code, DevOps, and SRE principles to manage networks programmatically.
  • Experience building workflows and pipelines to test and safely deploy production changes.
  • Platform engineering experience and the ability to build infrastructure for large-scale distributed systems.
  • Proficiency creating automation and tools with Python or Go.
  • Experience integrating LLMs, MCP, and agentic workflows into engineering processes.
  • Strong communication, documentation, collaboration, critical thinking, and self-directed problem-solving skills.

Tech Stack

Alibaba CloudAmbassadorAWSAzureDockerGoGrafanaKubernetesPrometheusPython

Categories

DevOpsSite Reliability