5 days ago
San Francisco, CA, USAMid Level
Base Salary
$155k - $200k/yr
Responsibilities
- Integrate internal and open-weight language and multimodal models into GPU evaluation and inference environments.
- Build automated benchmarks for model quality, latency, throughput, memory usage, and systems performance.
- Create standardized and reproducible comparisons across models, baselines, and runtime configurations.
- Build and maintain experiment tracking, model registries, and versioning for models, datasets, and evaluation configurations.
- Automate reproducible workflows from research checkpoints to validated deployments.
- Monitor model quality and systems performance and diagnose failures or regressions across model and deployment pipelines.
- Build reusable tools for researchers to launch evaluations, compare experiments, and reproduce results.
- Profile end-to-end model workloads and collaborate with distributed systems, inference, and GPU kernel engineers on performance issues.
Requirements
- At least 2 years of professional ML or software engineering experience, including production ML systems, ML platforms, or MLOps infrastructure.
- Strong Python and software engineering skills with experience building reliable production systems.
- Hands-on experience with PyTorch, TensorFlow, or JAX and understanding of modern language or multimodal model architectures.
- Experience with model evaluation or benchmarking and model lifecycle workflows such as experiment tracking, versioning, deployment, or monitoring.
- Experience running, benchmarking, and debugging models with GPU inference runtimes such as vLLM, SGLang, or TensorRT-LLM in containerized cloud or on-premises environments.
- Ability to document systems clearly and collaborate across research, infrastructure, and product engineering teams.
- MS or PhD in Computer Science, Computer Engineering, Machine Learning, or a related technical field, or equivalent practical experience.
- Familiarity with Hugging Face Transformers or similar model libraries is preferred.
- Experience enabling models on AMD GPUs and ROCm is preferred.
- Contributions to open-source evaluation, model, or ML infrastructure projects are preferred.
Benefits
- Medical, dental, and vision insurance.
- 401k plan.
- Daily lunch, snacks, and beverages.
- Flexible time off.
- Competitive salary and equity.
