Nace AI

Senior MLOps Engineer

Nace AI
Apply
2 months ago
Palo Alto, CA, USASenior

Responsibilities

  • Design, build, and operate end-to-end ML infrastructure for training orchestration, experiment tracking, model registries, automated evaluation, and model delivery.
  • Own LLM and SLM serving infrastructure, including low-latency, high-throughput inference, batching, caching, and autoscaling.
  • Build and manage multi-GPU training and inference clusters across cloud and on-premises environments while optimizing scheduling, utilization, and cost.
  • Implement production model observability for latency, throughput, drift, regression, and quality with actionable alerting.
  • Apply quantization, distillation support, KV-cache management, and other inference-time deployment optimizations.
  • Harden enterprise deployments through reproducibility, versioning, access controls, and audit-ready model traceability.
  • Establish MLOps best practices and infrastructure tooling standards as an early senior member of the infrastructure team.

Requirements

  • At least five years of experience in MLOps, ML infrastructure, or platform engineering with substantial production ownership.
  • Production experience deploying and scaling LLM inference infrastructure using serving frameworks such as TRT, vLLM, SGLang, or TGI.
  • Strong proficiency with Kubernetes, Docker, and infrastructure-as-code such as Terraform.
  • Hands-on experience managing GPU clusters and distributed training or serving environments.
  • Proficiency in Python and experience building substantial, maintainable systems.
  • Experience with ML pipeline and orchestration tools such as Airflow, Kubeflow, Ray, MLflow, or Weights & Biases.
  • Strong computer science fundamentals and cloud architecture experience with AWS, GCP, or Azure.
  • Bachelor’s degree in computer science or a related technical field.
  • Preferred: master’s degree, multi-node GPU training experience, quantization and inference optimization experience, Spark familiarity, LLM/VLM fine-tuning workflows, regulated or enterprise experience, or contributions to open-source ML infrastructure projects.

Benefits

  • Full-time, on-site role in Palo Alto, California
  • Significant equity and premium benefits
  • Opportunity to shape infrastructure foundations as an early infrastructure team member
Nace AI

About Nace AI

11-50 employees

Enterprise AI product & research company, building long-horizon reasoning models and agents. Our first product, Agentic Accounting, executes Financial Audit, Billing Audit, and Revenue Leakage Detection.

Contact me