SpreeAI

Principal Engineer, AI Platform & Infrastructure

SpreeAI
Apply
5 months ago

Responsibilities

  • Build and operate an end-to-end ML platform covering training, evaluation, deployment, and monitoring.
  • Create scalable training workflows with orchestration, infrastructure, resource management, model packaging, registries, dataset lineage, experiment tracking, checkpointing, and deployment automation.
  • Build reliable inference deployments and model delivery pipelines with versioning, reproducibility, rollback, approval gates, evaluation gates, and production observability.
  • Design and operate GPU allocation, scheduling, utilization, benchmarking, and resource-efficiency systems across training and inference workloads.
  • Establish production SLOs for latency, availability, error rate, GPU saturation, cold-start time, inference cost, and model quality drift.
  • Standardize inference serving infrastructure and optimize batching, quantization, CUDA graphs, memory-aware scheduling, throughput, latency, reliability, and cost.
  • Partner with research teams to productionize new multimodal and generative AI capabilities and drive execution across teams.

Requirements

  • 10+ years of software engineering or infrastructure experience, including 5+ years in ML infrastructure, MLOps, distributed systems, or AI platform engineering.
  • Deep experience with Python, PyTorch, Kubernetes, Docker, cloud infrastructure, and GPU-based workloads.
  • Strong understanding of distributed systems and large-scale ML infrastructure design.
  • Experience with ML workflow orchestration systems such as Ray, Kubeflow, Argo, Airflow, Flyte, or Metaflow.
  • Experience deploying and managing production inference systems using Triton, vLLM, TensorRT-LLM, Ray Serve, KServe, Seldon, BentoML, TorchServe, or custom services.
  • Experience with model registries, experiment tracking, CI/CD for ML, canary deployments, shadow traffic, rollback strategies, and production monitoring.
  • Strong cloud experience across AWS, GCP, Azure, CoreWeave, Lambda Labs, or RunPod.
  • Ability to debug performance bottlenecks across distributed systems, containers, networking, GPU memory, and storage layers.
  • Ability to define architecture, establish platform standards, and drive execution across teams.
  • Preferred: experience with multimodal, vision, generative AI, large-scale GPU clusters, A100/H100, NCCL, high-throughput data pipelines, generative AI evaluation and monitoring, ML security, privacy, data governance, or internal developer platforms.

Benefits

  • Build core AI infrastructure that directly shapes the deployment, monitoring, scalability, performance, and cost efficiency of real-world AI products.
  • Own critical infrastructure decisions end-to-end and work on high-leverage GPU efficiency and large-scale deployment problems.
  • Work in a high-velocity, low-bureaucracy environment with close collaboration across research, platform, and product teams.
  • Help accelerate AI capabilities into partner-facing fashion and e-commerce experiences.
SpreeAI

About SpreeAI

11-50 employees

SpreeAI builds AI-powered virtual try-on, sizing, and styling tools for fashion retailers and brands, delivered through web and mobile integrations and a partner portal. The company sells SaaS and SDKs that embed photorealistic try-on into ecommerce sites, in-store displays, and clienteling apps to reduce returns and improve conversion. Founded in 2023 and headquartered in Los Angeles, it integrates with platforms like Shopify to onboard brand partners quickly.

Contact me