
Staff Embedded ML Engineer, Edge AI
SimpliSafe2 months ago
Boston, MA, USAStaff+
Base Salary
$186k - $245k/yr
Responsibilities
- Own embedded deployment and performance optimization for real-time, on-device ML inference used in outdoor monitoring.
- Optimize latency, throughput, memory footprint, power, thermals, startup time, and stability across CPU, DSP, NPU, and GPU hardware.
- Perform kernel and operator optimization, including vectorization, tiling, cache-friendly layouts, memory-copy reduction, operator fusion, synchronization reduction, and thread scheduling.
- Integrate and maintain ML models in embedded pipelines using C/C++, including model validation, graph transforms, pre/post-processing, error handling, watchdogs, and fallback behavior.
- Drive INT8 and FP16 quantization readiness, calibration, numerical accuracy checks, and debugging of operator mismatches on target runtimes.
- Build profiling, benchmarking, memory-tracking, thermal-testing, and automated performance-regression tooling across device tiers and firmware versions.
- Partner with ML, firmware, and hardware teams to define deployment constraints, benchmarks, KPIs, and performance tradeoffs.
- Provide Staff-level leadership by setting standards, leading technical reviews, mentoring engineers, and influencing the on-device ML platform roadmap.
Requirements
- 8+ years of experience in embedded systems and/or performance engineering, including production software shipped on constrained devices.
- Strong C/C++ expertise and deep knowledge of CPU architecture, memory hierarchy, concurrency, and real-time considerations.
- Demonstrated experience optimizing embedded ML inference, including operator/kernel tuning and end-to-end pipeline optimization.
- Experience with modern vision model families such as DEIM, DFINE, RT-DETR, and YOLO or similar models.
- Experience with embedded inference runtimes and deployment workflows such as TFLite, ONNX Runtime, TensorRT, or vendor runtimes.
- Strong debugging and profiling skills using performance analysis, flame graphs, hardware counters, and tracing.
- Ability to lead cross-functionally across ML, firmware, and hardware teams and define benchmarks, KPIs, and technical tradeoffs.
- Preferred experience with embedded accelerators, DSP/NPU compilers, delegates, GPU compute, custom runtimes, and vendor toolchains.
- Preferred ARM NEON/SVE, hand-tuned kernel, XNNPACK, QNNPACK, oneDNN, or CMSIS-NN experience.
- Preferred large-scale INT8 quantized inference experience, including calibration, numerical debugging, overflow/underflow handling, and accuracy-performance tradeoffs.
- Preferred experience with camera or doorbell pipelines, ISP, video codecs, DMA, zero-copy buffers, and multithreaded real-time streaming.
- Preferred exposure to embedded Linux, RTOS, power management, thermal throttling, sustained-load performance, security/privacy for edge devices, and device-lab automation.
Benefits
- Hybrid work model with two core in-office days typically Tuesday, Wednesday, or Thursday and remaining workdays flexible between office and home.
- Comprehensive total rewards package with medical, retirement, lifestyle, wellness, and family-support benefits.
- Free SimpliSafe system and professional monitoring for the employee’s home.
- Employee Resource Groups offering networking, mentoring, development, and advocacy opportunities.
- Inclusive, mission- and values-driven workplace with a focus on growth, collaboration, and professional development.