
Staff Software Engineer, ML Infrastructure
SimpliSafe2 months ago
Boston, MA, USAStaff+
Base Salary
$147k - $215k/yr
Responsibilities
- Drive architecture and technical direction for the Kubernetes-based ML platform using Ray, KServe, Triton, and vLLM.
- Design and operate real-time cloud-side computer vision inference systems processing live video and device events at scale.
- Improve throughput, latency, GPU utilization, autoscaling, multi-model serving, reliability, and cost for production ML systems.
- Develop production LLM/GenAI serving infrastructure, including model-serving patterns, KV-cache and batching strategies, evaluation pipelines, guardrails, and cost controls.
- Lead technical reviews, capacity planning, incident response, postmortems, SLO definition, observability standards, and on-call practices.
- Establish model lifecycle management practices covering registries, deployment, monitoring, rollback, and drift.
- Mentor engineers through design reviews, code reviews, pairing, and written guidance, and create durable documentation and runbooks.
Requirements
- 8+ years of software engineering experience building and operating large-scale distributed systems in production.
- Deep expertise in high-throughput, low-latency services such as ad serving, recommendations, real-time APIs, or online platforms.
- Strong production experience with Kubernetes and AWS, including EKS, S3, IAM, and networking, plus Kafka, containerized deployments, and infrastructure-as-code.
- Experience with load balancing, autoscaling, batching, caching, multi-tenancy, queuing, and capacity planning.
- Proficiency in Python is required; Go, C++, or Rust experience for performance-sensitive components is preferred.
- Ability to lead ambiguous, cross-cutting technical initiatives, align senior stakeholders, and mentor engineers without formal authority.
- Strong written and verbal communication skills for explaining technical tradeoffs to ML scientists, product teams, and infrastructure teams.
- ML exposure, production ML systems experience, or ML-adjacent infrastructure experience is preferred but not required.
- Bonus qualifications include Ray, KServe, Triton, vLLM, TGI, TensorRT-LLM, SGLang, real-time video or streaming pipelines, GPU inference systems, ML lifecycle tooling, open-source distributed-systems contributions, and experience with strong security and compliance requirements.
Benefits
- Hybrid work model with two core in-office days, typically Tuesday, Wednesday, or Thursday, and flexibility to work from home the rest of the week.
- Comprehensive total rewards package with medical, retirement, wellness, lifestyle, and other benefits.
- Free SimpliSafe system and professional monitoring for the employee’s home.
- Employee Resource Groups offering networking, mentoring, development, and advocacy opportunities.
- Inclusive, mission- and values-driven culture with opportunities to build, grow, and thrive.