1 month ago
Base Salary
$165k - $330k/yr
Responsibilities
- Design and build scalable telemetry ingest and storage pipelines for metrics, logs, and traces across multi-cloud infrastructure.
- Own and evolve observability platforms through migrations and architectural improvements that improve reliability, reduce cost, and scale with organizational growth.
- Build instrumentation libraries, SDKs, and integrations that help engineering teams emit high-quality telemetry.
- Drive alerting and SLO infrastructure that enables teams to define, monitor, and respond to reliability targets with minimal noise.
- Partner with Inference, Product, and Infrastructure teams to meet their observability needs.
Requirements
- Deep experience with at least one observability signal area—metrics, logging, tracing, or error analytics—and familiarity with the others.
- Understanding of high-throughput data pipelines, columnar storage engines, and telemetry ingestion and querying tradeoffs at scale.
- Experience operating or building on observability platforms such as Prometheus, Grafana, ClickHouse, OpenTelemetry, or similar systems.
- Strong proficiency in at least one of Python, Rust, or Go.
- Strong communication skills and interest in partnering with internal teams on operational visibility and incident response.
- Interest in applying AI and LLMs to automated root cause analysis, anomaly detection, or intelligent alerting.
- Comfort working independently on ambiguous, high-impact infrastructure challenges.
Benefits
- 100% medical, dental, and vision insurance coverage for employees and dependents
- Flexible paid time off and company-wide Winter Break from Christmas Eve through New Year's Day
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k)
- Exposure to a variety of machine learning startups for learning and networking
About Baseten
Baseten builds an AI inference platform that provides tooling, infrastructure, and hardware to deploy, scale, and serve machine-learning models in production. The company sells managed model serving and developer tooling to software teams at AI product companies, with customers including Notion, Abridge, Writer, and Cursor. Privately held and headquartered in San Francisco, it focuses on high-availability, globally distributed inference for production workloads.
