Senior Machine Learning Engineer - Kubogent
Aivar Innovations Private Limited3 hours ago
Bengaluru, IndiaMid Level / Senior
Responsibilities
- Build reusable training and fine-tuning workflows and training-job abstractions for machine learning models and large language models.
- Develop inference workflows covering model loading, serving, batching, runtime configuration, evaluation, scaling, latency, and throughput.
- Perform data preparation, model training, fine-tuning, evaluation, debugging, and iteration.
- Implement reproducible MLOps pipelines for experimentation, evaluation, model registration, versioning, promotion, deployment, and rollback.
- Integrate experiment tracking and model lifecycle tooling, including metrics, artifacts, metadata, registries, and reproducibility.
- Define evaluation pipelines, quality metrics, benchmark datasets, regression checks, and model acceptance criteria.
- Optimize memory use, accelerator utilization, throughput, latency, batch sizing, precision, checkpointing, and parallelism.
- Make ML workflows production-grade through failure recovery, observability, artifact management, configuration, and reproducibility.
- Help shape Kubogent APIs, workflows, defaults, templates, and reusable platform primitives.
Requirements
- 4-8 years of relevant engineering experience with substantial hands-on work building and training machine learning systems.
- Hands-on experience training or fine-tuning large language models, including tokenization, dataset preparation, checkpoints, evaluation, and practical LLM constraints.
- Strong data handling skills covering cleaning, normalization, filtering, transformation, sampling, labeling, validation, and train/validation/test preparation.
- Strong Python skills and hands-on experience with mainstream ML frameworks such as PyTorch or TensorFlow.
- Experience with the modern LLM ecosystem, including Hugging Face Transformers and related model, tokenizer, dataset, and training tooling.
- Working experience with MLflow or comparable MLOps tooling, including experiment tracking, artifacts, parameters, metrics, model versioning, registries, and reproducible runs.
- Experience with model serving or inference runtimes and understanding of batching, latency, throughput, memory use, model loading, precision, and accelerator constraints.
- Ability to turn experimental ML code into maintainable libraries, services, jobs, pipelines, or platform components.
- Understanding of reproducibility across dataset, code, environment, configuration, seeds, checkpoints, artifacts, and metadata.
- Ability to debug across models, data, frameworks, runtimes, hardware, and surrounding systems.
- Preferred experience with distributed training technologies such as PyTorch Distributed, FSDP, DeepSpeed, or Ray Train.
- Preferred experience with parameter-efficient fine-tuning such as LoRA; LLM serving runtimes such as vLLM, NVIDIA Triton, TensorRT-LLM, or TGI; GPUs and ML accelerators; ML workloads on Kubernetes; ML workflow systems; large-dataset processing tools; model evaluation, benchmark, safety, or regression pipelines; or ML and developer platforms.