17 days ago
Remote, Germany +10 moreSenior
Responsibilities
- Build tools, automation, and workflows that simplify infrastructure-heavy tasks for AI teams.
- Develop monitoring, logging, and tracing systems for ML workflow performance and reproducibility.
- Design, implement, and maintain end-to-end machine learning pipelines for model and intelligent-agent development, training, and deployment.
- Work with large-scale distributed systems, including GPU clusters, to support model training, fine-tuning, and evaluation.
- Translate product and development goals into scalable and maintainable systems.
- Optimize ML workflows for reproducibility, scalability, cost efficiency, and team productivity.
Requirements
- At least three years of experience writing clean, maintainable Python code in modern ML codebases.
- Hands-on experience with MLOps tooling, Kubernetes, GCP, AWS, and ML orchestration frameworks.
- Understanding of the machine learning lifecycle from ideation through customer-facing applications.
- Ability to own projects end to end, including design, experimentation, implementation, and iteration.
- Experience with CI/CD systems such as GitHub Actions or JetBrains TeamCity.
- Preferred experience with ZenML, Dagster, Airflow, Kubernetes-based infrastructure, Python backend services, ML pipelines, and experiment tracking or observability tools.
- Especially valued experience with vLLM, DeepSpeed, TensorRT, Python libraries for ML engineers, NLP, transformer-based approaches, Java, or Kotlin.
Benefits
- Hybrid work arrangement, indicated by the #LI-HYBRID designation.
- Equal opportunity and inclusive workplace.
Tech Stack
Categories
About JetBrains
On a mission to make software development a more productive and enjoyable experience. Make it happen. With Code.
