
Senior AI Infra JD for Inference & Eval's
Ambient.aiBase Salary
$168k - $205k/yr
Responsibilities
- Design, build, and maintain AI infrastructure for real-time computer vision, LLM, LVM, and multimodal inference workloads.
- Build scalable systems for running models across large volumes of video and sensor data.
- Optimize inference latency, throughput, GPU utilization, reliability, and cost.
- Develop evaluation harnesses and benchmarking systems for model quality, system performance, regressions, and production readiness.
- Build infrastructure for continuous model evaluation, experimentation, and deployment.
- Partner with research scientists to productionize advances in computer vision, LLMs, LVMs, RAG, and multimodal AI.
- Improve model-serving architecture through batching, caching, routing, quantization, model parallelism, and hardware utilization.
- Develop data engines and feedback loops for training-data collection, model-behavior evaluation, and continuous AI improvement.
- Create observability, monitoring, and debugging tools for production AI systems.
- Define best practices for deploying, evaluating, and operating AI systems in enterprise environments.
Requirements
- 4+ years of industry experience building infrastructure, distributed systems, machine learning platforms, or production AI systems.
- BS/MS in Computer Science or a related technical field, or equivalent practical experience.
- Strong programming skills, especially in Python, with solid software engineering fundamentals.
- Experience designing and building scalable machine learning infrastructure for training, inference, evaluation, and deployment.
- Hands-on experience running deep learning models in production, ideally including LLMs, LVMs, vision-language models, or multimodal models.
- Strong understanding of batching, caching, quantization, parallelism, memory optimization, GPU utilization, and latency reduction.
- Experience with model-serving frameworks such as vLLM or Triton Inference Server.
- Experience building evaluation frameworks, test harnesses, benchmarks, regression tests, or model-quality measurement systems.
- Strong machine learning and deep learning background; computer vision experience is a strong plus.
- Experience designing data engines or pipelines for collecting, managing, and curating training and evaluation data.
- Familiarity with integrating LLMs, LVMs, RAG pipelines, embedding models, or multimodal models into production applications.
- Experience with cloud infrastructure, containers, orchestration, distributed systems, and GPU-based workloads.
- Strong collaboration and communication skills across research, product, infrastructure, and stakeholder groups.
- Experience operating large-scale GPU infrastructure or distributed inference systems is preferred.
- Experience with CUDA, NCCL, PyTorch, TensorRT, or ONNX is preferred.
- Experience with video understanding, real-time computer vision, multimodal AI, or physical-world AI systems is preferred.
- Experience with model compression, speculative decoding, distillation, pruning, or low-latency serving is preferred.
- Experience with prompt evaluation, model regression testing, human-in-the-loop evaluation, or automated quality gates is preferred.
- Familiarity with retrieval-augmented generation, vector databases, embedding models, re-rankers, or search infrastructure is preferred.
- Experience building internal ML platforms or tools for researchers and applied ML teams is preferred.
Benefits
- Regular full-time employees receive stock options.
- Comprehensive health and welfare benefits include medical, dental, vision, life, EAP, legal services, and a 401(k) plan.
- Flexible time off includes a Winter Break for most roles, depending on customer demand.
- The latest technology and company swag are delivered to employees' doors.
- Redwood City engineering, product, design, and marketing teams work from the office three days per week; other Bay Area employees work in the office on Fridays.
- Opportunities to connect with coworkers and participate in company social activities are available.
Categories
About Ambient.ai
Ambient.ai is the leader in Agentic Physical Security. At the core of the platform is Ambient Intelligence, powered by Ambient Pulsar, the first always-on, edge-optimized reasoning Vision-Language Model (VLM) purpose-built for physical security. Ambient.ai transforms existing cameras, sensors, and access control systems into a unified intelligence layer that continuously perceives, understands, and acts in real time. By reasoning across video, access, and sensor data, Ambient.ai detects and interprets 150+ threat signatures, validates real risk, and orchestrates the appropriate response. With capabilities spanning Agentic Monitoring, Agentic Investigations, Agentic Access Intelligence, and Agentic Threat Analysis & Response, customers reduce false alarms by over 90%, accelerate investigations from days to minutes, and help resolve more than 80% of alerts in under one minute. Trusted by Fortune 100 enterprises across campuses, data centers, and critical infrastructure, Ambient.ai delivers proactive safety, operational efficiency, and measurable ROI at enterprise scale.