
Staff Machine Learning Engineer, ML Infrastructure
SimpliSafe3 months ago
Boston, MA, USAStaff+
Base Salary
$184k - $245k/yr
Responsibilities
- Drive architecture and technical direction for a Kubernetes-based ML platform using Ray, KServe, Triton, and vLLM across real-time and batch workloads.
- Design and evolve cloud-side real-time computer vision inference systems processing live video and device events.
- Improve inference throughput, latency, GPU utilization, autoscaling, multi-model serving, reliability, and cost.
- Develop production LLM/GenAI serving patterns, evaluation pipelines, guardrails, KV-cache and batching strategies, and cost controls.
- Establish model lifecycle management, deployment, monitoring, rollback, drift, observability, SLO, and on-call practices.
- Lead incident response and postmortems for critical ML systems and convert findings into platform improvements.
- Mentor engineers through design reviews, code reviews, pairing, and written guidance while documenting platform standards and architecture.
- Partner with applied ML engineers to move GenAI product features from prototype to scaled deployment.
Requirements
- 8+ years of software or ML engineering experience building and operating production ML systems at scale.
- Deep expertise in cloud ML infrastructure on Kubernetes and hands-on production experience with Ray.
- Strong production experience with AWS, including EKS, S3, IAM, and networking, plus Kafka, containerized deployments, CI/CD, and infrastructure-as-code.
- Experience designing and operating high-throughput, low-latency inference systems with GPU-aware scheduling, batching, autoscaling, and multi-tenancy.
- Solid understanding of model training, evaluation, versioning, deployment, monitoring, and rollback in production.
- Proficiency in Python; experience with Go, C++, or Rust for performance-sensitive components is a plus.
- Staff-level technical leadership, including driving ambiguous cross-cutting initiatives, aligning stakeholders, and mentoring without formal authority.
- Strong written and verbal communication skills for explaining technical tradeoffs to ML scientists, product teams, and infrastructure teams.
- Preferred experience includes production LLM serving with vLLM, TGI, TensorRT-LLM, or SGLang; real-time video or streaming ML pipelines; computer vision workloads; model lifecycle tooling; open-source ML infrastructure contributions; and security or compliance environments.
Benefits
- Hybrid work model with two core in-office days, typically Tuesday, Wednesday, or Thursday, and flexibility to work from home for the remainder of the week.
- Comprehensive total rewards package with medical, retirement, lifestyle, wellness, and family-support benefits.
- Free SimpliSafe system and professional monitoring for the employee’s home.
- Employee Resource Groups offering networking, mentoring, development, and advocacy opportunities.
- Inclusive, mission-driven culture with opportunities to grow and make an impact on home security.