Sr Site Reliability Engineer
The Trade Desk5 months ago
Responsibilities
- Design, build, and scale a global network platform across physical datacenters and AWS, Azure, and Alibaba Cloud.
- Support thousands of hosts worldwide and engineer reliable solutions for petabyte-scale data challenges.
- Troubleshoot and resolve complex network issues while maintaining high availability and performance.
- Lead root cause analyses and postmortems and convert incidents into operational improvements.
- Build tools, automate workflows, and eliminate operational toil.
- Participate in a global follow-the-sun on-call rotation.
- Partner with SRE and infrastructure teams to shape network automation strategy and build scalable, maintainable solutions.
Requirements
- 6–8 years of hands-on network automation and operational experience supporting large-scale production infrastructure.
- Strong software development and networking experience with a software-first mindset.
- Deep expertise in TCP/IP, the OSI model, BGP, OSPF, and large-scale IP networking.
- Hands-on experience with Kubernetes networking technologies such as Cilium and Calico and understanding of CNIs.
- Experience managing NGINX Ingress, Envoy, or HAProxy in large-scale production environments.
- Experience troubleshooting and performance tuning Kubernetes and Docker networking; bare-metal Kubernetes experience is a plus.
- Knowledge of IPv6, SDN, SDN controllers, QoS, and bandwidth management.
- Experience operating SONiC, Cisco IOS, JunOS, Arista EOS, or Nokia SR Linux/SR OS at scale.
- Experience with monitoring, alerting, complex rules, and time-series queries using tools such as Prometheus and Grafana.
- Experience applying infrastructure-as-code, DevOps, and SRE principles to manage networks programmatically.
- Experience building workflows and pipelines to test and safely deploy production changes.
- Platform engineering experience and the ability to build infrastructure for large-scale distributed systems.
- Proficiency creating automation and tools with Python or Go.
- Experience integrating LLMs, MCP, and agentic workflows into engineering processes.
- Strong communication, documentation, collaboration, critical thinking, and self-directed problem-solving skills.