Parallel Systems

Senior ML Ops Engineer (Machine Learning Infrastructure)

Parallel Systems
Apply
2 months ago

Base Salary

$150k - $250k/yr

Responsibilities

  • Design and implement automated MLOps pipelines for data management, model training, deployment, and monitoring.
  • Architect, deploy, and manage scalable infrastructure for distributed ML training and inference.
  • Collaborate with ML engineers to define data management, model development, and deployment strategies.
  • Build and operate AWS- or GCP-based cloud systems optimized for ML workloads in R&D and production.
  • Develop infrastructure supporting continuous integration and deployment, experiment management, and model and dataset governance.
  • Automate model evaluation, selection, and deployment workflows.
  • Own the end-to-end ML infrastructure stack and deliver its core features, including integrations with tools such as MLflow, SageMaker, or Kubeflow.

Requirements

  • Bachelor’s or higher degree in Computer Science, Machine Learning, or a relevant engineering discipline.
  • At least 5 years of experience building large-scale, reliable systems, including at least 2 years focused on ML infrastructure or MLOps.
  • Experience architecting and deploying production-grade ML pipelines and platforms across the ML lifecycle.
  • Hands-on experience with MLOps tools such as MLflow, Kubeflow, SageMaker, Airflow, or Metaflow.
  • Strong understanding of CI/CD practices applied to ML workflows and proficiency in Python, Git, and system design.
  • Experience designing ML architectures on AWS, GCP, or Azure.
  • Preferred experience with deep learning architectures such as CNNs, RNNs, and Transformers or with computer vision.
  • Preferred experience with distributed training tools such as PyTorch DDP, Horovod, or Ray.
  • Preferred background in real-time ML systems, batch inference, CPU/GPU-aware orchestration, autonomous vehicles, robotics, or other real-time ML-driven systems.

Benefits

  • Hybrid role with a minimum of one week per month onsite in Los Angeles.

Tech Stack

Parallel Systems

About Parallel Systems

51-200 employees

Parallel Systems builds autonomous, battery-electric rail vehicles and fleet-management software that move containerized freight by rail, serving railroads and shippers seeking to shift short-haul loads off trucks. Privately held and founded in 2020 in Los Angeles, the company integrates with existing rail operations and sells vehicles and software-as-a-service to U.S. freight rail operators and logistics providers.

Contact me