
Tech Lead – Network Observability
Clockwork.io4 months ago
Base Salary
$180k - $260k/yr
Responsibilities
- Lead the architecture, design, and development of scalable network monitoring platforms for RDMA, RoCE, InfiniBand, and TCP/IP infrastructure
- Build backend telemetry services, observability dashboards, alerts, diagnostics, anomaly detection, SLA monitoring, and traffic analysis workflows
- Troubleshoot production issues across application, operating system, server, RDMA, and network layers while optimizing low-latency collection, aggregation, and alerting
- Establish engineering standards, drive automation, define technical roadmaps with cross-functional teams, and mentor engineers on distributed systems and high-performance networking
Requirements
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field
- Hands-on programming experience in C++, Go, Python, Rust, or similar systems programming languages
- Experience leading engineering teams, major technical initiatives, or complex infrastructure projects
- Experience building distributed systems, backend services, telemetry pipelines, or observability platforms
- Hands-on experience with RDMA, RoCE, InfiniBand, or other high-performance network fabrics
- Familiarity with libibverbs, RDMA verbs, RDMA CM, queue pairs, completion queues, and memory registration
- Strong knowledge of Linux networking, TCP/IP, DNS, HTTP, routing, MTU, congestion control, packet loss, and performance tuning
- Experience with traceroute-style diagnostics, path discovery, network reachability checks, synthetic probes, or active network measurements
- Experience with monitoring and visualization platforms such as Prometheus, Grafana, Datadog, Splunk, or OpenTelemetry
- Strong debugging skills across software, operating system, server, and network layers
- Experience operating production systems in Linux-based environments
- Strong architectural judgment and ability to design reliable, scalable, and operationally simple systems
- Experience supporting AI/ML, HPC, storage, or GPU cluster infrastructure workloads is preferred
- Experience with large-scale RoCE or InfiniBand deployments, NCCL, distributed training infrastructure, or AI cluster diagnostics is preferred
- Experience with eBPF, XDP, DPDK, perf, tcpdump, Wireshark, ethtool, iproute2, rdma-core, or Linux kernel networking tools is preferred
- Experience with AWS, GCP, or Azure is preferred
- Experience with Kubernetes, service discovery, configuration management, and infrastructure automation is preferred
- Knowledge of security, compliance, and infrastructure best practices is preferred
- Experience designing time-series data systems, alerting pipelines, or high-cardinality telemetry platforms is preferred
Benefits
- Competitive compensation and a great benefits package
- Catered lunch
- Stock options or other equity awards may be available through the company’s equity program
- Friendly and inclusive workplace culture
- Challenging projects
Tech Stack
About Clockwork.io
Clockwork.io pioneers Software-Driven AI Fabrics™, delivering a programmable software layer that makes large-scale AI clusters observable, deterministic, and resilient by design to drive continuous workload progress and peak cluster utilization. Its FleetIQ platform enables enterprises to train, deploy, and serve the world's most demanding AI workloads faster, more reliably, and at lower cost. Companies including Uber, Wells Fargo, DCAI, Nebius, Nscale, and White Fiber trust Clockwork.io to power their AI infrastructure. Learn more at www.clockwork.io