about 3 hours ago
Base Salary
$165k - $330k/yr
Responsibilities
- Design and build scalable telemetry ingestion and storage pipelines for metrics, logs, and traces across multi-cloud infrastructure.
- Own and evolve observability platforms through migrations and architectural improvements that increase reliability, reduce cost, and support organizational growth.
- Build instrumentation libraries, SDKs, and integrations that help engineering teams emit high-quality telemetry.
- Develop alerting and SLO infrastructure that enables teams to define, monitor, and respond to reliability targets with minimal noise.
- Partner with Inference, Product, and Infrastructure teams to meet their operational visibility and reliability needs.
Requirements
- Deep experience with at least one observability signal area, such as metrics, logging, tracing, or error analytics, plus familiarity with the others.
- Understanding of high-throughput data pipelines, columnar storage engines, and telemetry ingestion and querying tradeoffs at scale.
- Experience operating or building on observability platforms such as Prometheus, Grafana, ClickHouse, or OpenTelemetry.
- Strong proficiency in at least one of Python, Rust, or Go.
- Strong communication and cross-functional partnership skills.
- Interest in applying AI and large language models to root cause analysis, anomaly detection, or intelligent alerting.
- Ability to work independently on ambiguous, high-impact infrastructure challenges.
Benefits
- Competitive compensation including meaningful equity.
- 100% coverage of medical, dental, and vision insurance for employees and dependents.
- Flexible paid time off, including a company-wide Winter Break from Christmas Eve through New Year's Day.
- Paid parental leave.
- Fertility and family-building stipend through Carrot.
- Company-facilitated 401(k).
- Exposure to a variety of ML startups and related learning and networking opportunities.
About Baseten
Inference is everything. Baseten is an AI infrastructure platform giving you the tooling, expertise, and hardware needed to bring great AI products to market - fast. Our proprietary Inference Stack utilizes the cutting-edge of performance research combined with highly performant and reliable infrastructure to give you out-of-the-box global availability with 99.99% of uptime.
