5 months ago
Base Salary
$180k - $440k/yr
Responsibilities
- Design and implement scalable infrastructure for metrics, logging, and tracing.
- Build high-performance telemetry pipelines capable of handling massive ingestion volumes.
- Develop APIs, query engines, and UIs that provide real-time service insights.
- Define and enforce company-wide best practices for instrumentation, alerting, and reliability.
- Integrate observability deeply with infrastructure, product teams, and internal platforms.
- Own the reliability, scalability, and performance of the observability stack end to end.
Requirements
- Production-level proficiency in Go, Rust, Scala, or a similar language.
- Deep understanding of distributed systems and telemetry architecture.
- Experience building and operating infrastructure at scale.
- Familiarity with Prometheus, Grafana, OpenTelemetry, VictoriaMetrics, or ClickHouse.
- Experience with Kafka, Redis, or large-scale time-series databases.
- Experience operating observability pipelines in Kubernetes or similar orchestration environments.
Benefits
- Equity and comprehensive medical, vision, and dental coverage.
- Access to a 401(k) retirement plan, short- and long-term disability insurance, life insurance, discounts, and other perks.
Tech Stack
About xAI
Understand the Universe. We are a team of AI technologists and business leaders on a mission to build AI systems that can help humanity understand the world better. https://x.ai/careers