2 months ago
Palo Alto, CA, USASenior
Responsibilities
- Design, build, and operate end-to-end ML infrastructure for training orchestration, experiment tracking, model registries, automated evaluation, and model delivery.
- Own LLM and SLM serving infrastructure, including low-latency, high-throughput inference, batching, caching, and autoscaling.
- Build and manage multi-GPU training and inference clusters across cloud and on-premises environments while optimizing scheduling, utilization, and cost.
- Implement production model observability for latency, throughput, drift, regression, and quality with actionable alerting.
- Apply quantization, distillation support, KV-cache management, and other inference-time deployment optimizations.
- Harden enterprise deployments through reproducibility, versioning, access controls, and audit-ready model traceability.
- Establish MLOps best practices and infrastructure tooling standards as an early senior member of the infrastructure team.
Requirements
- At least five years of experience in MLOps, ML infrastructure, or platform engineering with substantial production ownership.
- Production experience deploying and scaling LLM inference infrastructure using serving frameworks such as TRT, vLLM, SGLang, or TGI.
- Strong proficiency with Kubernetes, Docker, and infrastructure-as-code such as Terraform.
- Hands-on experience managing GPU clusters and distributed training or serving environments.
- Proficiency in Python and experience building substantial, maintainable systems.
- Experience with ML pipeline and orchestration tools such as Airflow, Kubeflow, Ray, MLflow, or Weights & Biases.
- Strong computer science fundamentals and cloud architecture experience with AWS, GCP, or Azure.
- Bachelor’s degree in computer science or a related technical field.
- Preferred: master’s degree, multi-node GPU training experience, quantization and inference optimization experience, Spark familiarity, LLM/VLM fine-tuning workflows, regulated or enterprise experience, or contributions to open-source ML infrastructure projects.
Benefits
- Full-time, on-site role in Palo Alto, California
- Significant equity and premium benefits
- Opportunity to shape infrastructure foundations as an early infrastructure team member
Tech Stack
Categories
About Nace AI
Enterprise AI product & research company, building long-horizon reasoning models and agents. Our first product, Agentic Accounting, executes Financial Audit, Billing Audit, and Revenue Leakage Detection.
