Aivar Innovations Private Limited

Senior Machine Learning Engineer - Kubogent

Aivar Innovations Private Limited
Apply
3 hours ago
Bengaluru, IndiaMid Level / Senior

Responsibilities

  • Build reusable training and fine-tuning workflows and training-job abstractions for machine learning models and large language models.
  • Develop inference workflows covering model loading, serving, batching, runtime configuration, evaluation, scaling, latency, and throughput.
  • Perform data preparation, model training, fine-tuning, evaluation, debugging, and iteration.
  • Implement reproducible MLOps pipelines for experimentation, evaluation, model registration, versioning, promotion, deployment, and rollback.
  • Integrate experiment tracking and model lifecycle tooling, including metrics, artifacts, metadata, registries, and reproducibility.
  • Define evaluation pipelines, quality metrics, benchmark datasets, regression checks, and model acceptance criteria.
  • Optimize memory use, accelerator utilization, throughput, latency, batch sizing, precision, checkpointing, and parallelism.
  • Make ML workflows production-grade through failure recovery, observability, artifact management, configuration, and reproducibility.
  • Help shape Kubogent APIs, workflows, defaults, templates, and reusable platform primitives.

Requirements

  • 4-8 years of relevant engineering experience with substantial hands-on work building and training machine learning systems.
  • Hands-on experience training or fine-tuning large language models, including tokenization, dataset preparation, checkpoints, evaluation, and practical LLM constraints.
  • Strong data handling skills covering cleaning, normalization, filtering, transformation, sampling, labeling, validation, and train/validation/test preparation.
  • Strong Python skills and hands-on experience with mainstream ML frameworks such as PyTorch or TensorFlow.
  • Experience with the modern LLM ecosystem, including Hugging Face Transformers and related model, tokenizer, dataset, and training tooling.
  • Working experience with MLflow or comparable MLOps tooling, including experiment tracking, artifacts, parameters, metrics, model versioning, registries, and reproducible runs.
  • Experience with model serving or inference runtimes and understanding of batching, latency, throughput, memory use, model loading, precision, and accelerator constraints.
  • Ability to turn experimental ML code into maintainable libraries, services, jobs, pipelines, or platform components.
  • Understanding of reproducibility across dataset, code, environment, configuration, seeds, checkpoints, artifacts, and metadata.
  • Ability to debug across models, data, frameworks, runtimes, hardware, and surrounding systems.
  • Preferred experience with distributed training technologies such as PyTorch Distributed, FSDP, DeepSpeed, or Ray Train.
  • Preferred experience with parameter-efficient fine-tuning such as LoRA; LLM serving runtimes such as vLLM, NVIDIA Triton, TensorRT-LLM, or TGI; GPUs and ML accelerators; ML workloads on Kubernetes; ML workflow systems; large-dataset processing tools; model evaluation, benchmark, safety, or regression pipelines; or ML and developer platforms.

Tech Stack

Apache SparkAWSHugging Face TransformersKubernetesMLflowPandasPythonPyTorchTensorFlow
Aivar Innovations Private Limited

About Aivar Innovations Private Limited

51-200 employees
Contact me