4 days ago
San Francisco, CA, USA or New York, NY, USASenior
Responsibilities
- Design, build, and operate core platform systems used daily by ML researchers and product engineers.
- Partner with research teams to turn recurring pain points into durable infrastructure.
- Own reliability, performance, and developer experience for assigned systems.
- Build telemetry, ML data platform, observability, or ML developer-experience and systems infrastructure.
- Ship iteratively, measure impact, and improve platform quality in a high-ownership environment.
Requirements
- Strong background in systems or infrastructure software engineering and building platforms used by other engineers.
- Experience owning production distributed systems at meaningful scale, such as ingestion systems, data pipelines, scheduling, or orchestration.
- Comfort working across Linux, cloud or bare-metal environments, and modern orchestration technologies such as Kubernetes or Ray.
- Ability to work closely with ML researchers and product engineers.
- Especially relevant experience includes event ingestion, product analytics pipelines, OpenTelemetry, tracing, data frameworks, Spark, Flink, Ray, ML datasets, training-data infrastructure, experiment monitoring, evaluation tooling, GPU or cluster scheduling, job queues, and node health.
Benefits
- In-person role with offices in North Beach, San Francisco; Palo Alto; and Manhattan, New York.
- Technical interview process includes 2–3 short technical interviews followed by an onsite small project, discussion, and team meetings.
- Offices include well-stocked libraries.
