4 months ago
Responsibilities
- Extend profiling and tracing tools for new accelerators, including performance-data collection, compression, and visualization.
- Develop CLI tools and automation for launching inference stacks, benchmarks, and configuration management.
- Convert internal tooling prototypes into high-performance, scalable, accessible commands.
- Build request tracing, pipeline visualization, and debugging tools for multi-agent serving workflows and complex inference DAGs.
- Create shared internal libraries and abstractions that improve engineering productivity.
Requirements
- Demonstrate strong software engineering fundamentals, including clean APIs, robust error handling, sensible defaults, and clear documentation.
- Have experience with profiling and tracing systems such as perf, Nsight, Tracy, or similar tools.
- Be familiar with observability stacks such as Prometheus, Grafana, OpenTelemetry, or equivalent systems.
- Be comfortable working across low-level trace collection, dashboards, and developer-facing CLI tools.
Benefits
- Competitive salary determined by skills and experience.
- Equity and ownership.
- Private healthcare.
- Visa sponsorship and relocation benefits.
- In-person work at the London office with provided tools, workspace, and setup.
Tech Stack
GrafanaPrometheus
