
Senior Software Engineer – Network Observability
Clockwork.io25 days ago
Base Salary
$140k - $210k/yr
Responsibilities
- Design, develop, and scale high-performance monitoring platforms for RDMA, RoCE, InfiniBand, and TCP/IP infrastructure.
- Build backend telemetry services, observability dashboards, alerts, diagnostics, anomaly detection, SLA monitoring, and traffic analysis workflows.
- Troubleshoot production issues across application, operating system, server, RDMA, and network layers.
- Optimize low-latency collection, aggregation, and alerting systems.
- Collaborate cross-functionally, improve engineering practices, automate operational workflows, and contribute to platform technical direction.
Requirements
- Bachelor’s or master’s degree in computer science, computer engineering, electrical engineering, or a related technical field.
- Strong hands-on programming experience in C++, Go, Python, Rust, or similar systems programming languages.
- Experience delivering complex infrastructure projects and designing and implementing distributed systems.
- Experience building distributed systems, backend services, telemetry pipelines, or observability platforms.
- Hands-on experience with RDMA, RoCE, InfiniBand, or other high-performance network fabrics.
- Familiarity with libibverbs, RDMA verbs, RDMA CM, queue pairs, completion queues, memory registration, and related RDMA concepts.
- Strong knowledge of Linux networking, TCP/IP, DNS, HTTP, routing, MTU, congestion control, packet loss, latency, and performance tuning.
- Experience with network diagnostics, path discovery, reachability checks, synthetic probes, or active network measurements.
- Experience with Prometheus, Grafana, Datadog, Splunk, OpenTelemetry, or similar monitoring and visualization tools.
- Strong debugging skills across software, operating system, server, and network layers.
- Experience operating production systems in Linux-based environments.
- Preferred experience includes AI/ML, HPC, storage, GPU cluster, large-scale RoCE or InfiniBand, NCCL, distributed training, eBPF, XDP, DPDK, perf, tcpdump, Wireshark, ethtool, iproute2, rdma-core, cloud infrastructure, Kubernetes, service discovery, configuration management, infrastructure automation, security, compliance, and time-series telemetry systems.
Benefits
- Competitive compensation and a great benefits package.
- Catered lunch.
- Friendly and inclusive workplace culture.
- Challenging projects.
- Additional equity participation may include stock options or other equity awards, subject to the company’s equity program and applicable approvals.
Tech Stack
Categories
About Clockwork.io
Clockwork.io pioneers Software-Driven AI Fabrics™, delivering a programmable software layer that makes large-scale AI clusters observable, deterministic, and resilient by design to drive continuous workload progress and peak cluster utilization. Its FleetIQ platform enables enterprises to train, deploy, and serve the world's most demanding AI workloads faster, more reliably, and at lower cost. Companies including Uber, Wells Fargo, DCAI, Nebius, Nscale, and White Fiber trust Clockwork.io to power their AI infrastructure. Learn more at www.clockwork.io