7 months ago
San Francisco, CA, USAMid Level
Responsibilities
- Design and build scalable infrastructure for AI model training and inference workflows.
- Develop high-performance APIs, backend services, and model-serving pipelines.
- Build distributed systems supporting large-scale generative and multimodal models.
- Optimize GPU utilization, latency, throughput, performance, and cost.
- Improve observability, monitoring, reliability, and production readiness of AI systems.
- Partner with Applied Science teams to productionize research systems.
- Drive deployment automation, workflow improvements, and developer-platform usability.
Requirements
- Degree in Computer Science, Engineering, or comparable education and practical experience.
- Strong object-oriented programming skills in Python, C++, Java, Go, or similar languages.
- Strong foundations in data structures and algorithms.
- Experience building production backend or distributed systems.
- Understanding of cloud infrastructure concepts and containerized systems.
- Preferred: experience with Kubernetes, Docker, or container orchestration.
- Preferred: familiarity with GPU-based machine-learning workloads or distributed training and inference systems.
- Preferred: experience with vLLM, Triton, Ray Serve, or similar model-serving frameworks.
- Preferred: experience with observability tools and performance debugging.
- Preferred: familiarity with PyTorch or machine-learning workflows.
Categories
About SpreeAI
SpreeAI builds AI-powered virtual try-on, sizing, and styling tools for fashion retailers and brands, delivered through web and mobile integrations and a partner portal. The company sells SaaS and SDKs that embed photorealistic try-on into ecommerce sites, in-store displays, and clienteling apps to reduce returns and improve conversion. Founded in 2023 and headquartered in Los Angeles, it integrates with platforms like Shopify to onboard brand partners quickly.
