6 months ago
Base Salary
$200k - $358k/yr
Responsibilities
- Design, build, and operate the end-to-end ML platform for training, experimentation, batch and online inference, and edge deployment.
- Partner with product and applied ML teams to launch and iterate ML-powered features, including computer vision and LLM-based reporting capabilities.
- Lead throughput, cost estimation, and capacity planning for ML features from exploration through production.
- Define experiment success metrics, structure A/B tests and offline evaluations, and interpret results for product and technical decisions.
- Evolve shared training and experimentation infrastructure, including orchestration, cluster configuration, environment management, experiment tracking, evaluation, and regression testing.
- Design and operate scalable Ray- and Spark-based inference systems with deployment patterns, observability, and SLOs.
- Partner with firmware and edge teams to package, validate, and deploy models to devices and build edge-to-cloud feedback loops.
- Own the reliability, observability, security posture, on-call practices, incident response, and infrastructure hardening of ML systems.
- Set ML infrastructure architecture and strategy, influence cross-team decisions, mentor engineers and applied scientists, and drive developer experience through documentation and office hours.
- Own or co-own high-priority technical delivery from modeling and system design through production rollout.
- Contribute to and represent Samsara in open source communities such as Ray, Spark, and RayDP.
Requirements
- 10+ years of experience in machine learning engineering or related fields, with a strong track record building and operating large-scale ML systems.
- Strong experience with distributed computing frameworks such as Ray and/or Spark.
- Hands-on experience with AWS, containers/Kubernetes, and production observability tooling.
- Proven experience building or supporting training, experimentation, or inference platforms used by multiple teams.
- Solid understanding of ML fundamentals, including evaluation, experiment design, and production model iteration.
- Experience shipping ML-powered features end-to-end with measurable product or business impact is preferred.
- Computer vision and/or LLM-based production experience is preferred.
- Experience with edge or on-device ML and collaboration with firmware or embedded teams is preferred.
- Familiarity with model registries, deployment, monitoring, rollback, and drift detection is preferred.
- Experience in environments with strong security and compliance requirements is preferred.
- Demonstrated ability to lead across teams and influence technical direction at Staff+ scope is preferred.
Benefits
- Remote position open to candidates based in the United States.
- Flexible employee-led remote working model with offices available for in-person work.
- Professional development stipend.
- Comprehensive health and parental leave plans.
- Initial RSU grant and ongoing refresh opportunities tied to performance, subject to plan terms and conditions.
Tech Stack
Categories
About Samsara
Samsara builds a Connected Operations Cloud that uses IoT hardware and software to help fleets and other physical-operations organizations monitor vehicles, assets, and worksites. Its platform spans video-based driver safety, vehicle telematics, equipment monitoring, and mobile workflows, sold as subscriptions with accompanying devices. Founded in 2015 and headquartered in San Francisco, Samsara is a public company on the NYSE (ticker: IOT) serving transportation, construction, manufacturing, field services, and government.
