
Senior ML Ops Engineer (Machine Learning Infrastructure)
Parallel Systems2 months ago
Base Salary
$150k - $250k/yr
Responsibilities
- Design and implement automated MLOps pipelines for data management, model training, deployment, and monitoring.
- Architect, deploy, and manage scalable infrastructure for distributed ML training and inference.
- Collaborate with ML engineers to define data management, model development, and deployment strategies.
- Build and operate AWS- or GCP-based cloud systems optimized for ML workloads in R&D and production.
- Develop infrastructure supporting continuous integration and deployment, experiment management, and model and dataset governance.
- Automate model evaluation, selection, and deployment workflows.
- Own the end-to-end ML infrastructure stack and deliver its core features, including integrations with tools such as MLflow, SageMaker, or Kubeflow.
Requirements
- Bachelor’s or higher degree in Computer Science, Machine Learning, or a relevant engineering discipline.
- At least 5 years of experience building large-scale, reliable systems, including at least 2 years focused on ML infrastructure or MLOps.
- Experience architecting and deploying production-grade ML pipelines and platforms across the ML lifecycle.
- Hands-on experience with MLOps tools such as MLflow, Kubeflow, SageMaker, Airflow, or Metaflow.
- Strong understanding of CI/CD practices applied to ML workflows and proficiency in Python, Git, and system design.
- Experience designing ML architectures on AWS, GCP, or Azure.
- Preferred experience with deep learning architectures such as CNNs, RNNs, and Transformers or with computer vision.
- Preferred experience with distributed training tools such as PyTorch DDP, Horovod, or Ray.
- Preferred background in real-time ML systems, batch inference, CPU/GPU-aware orchestration, autonomous vehicles, robotics, or other real-time ML-driven systems.
Benefits
- Hybrid role with a minimum of one week per month onsite in Los Angeles.
Tech Stack
Categories
About Parallel Systems
Parallel Systems builds autonomous, battery-electric rail vehicles and fleet-management software that move containerized freight by rail, serving railroads and shippers seeking to shift short-haul loads off trucks. Privately held and founded in 2020 in Los Angeles, the company integrates with existing rail operations and sells vehicles and software-as-a-service to U.S. freight rail operators and logistics providers.