1 month ago
Remote, IndiaSenior
Responsibilities
- Fine-tune and train small language models using Hugging Face, TRL, LoRA, QLoRA, and PEFT.
- Optimize models for inference using quantization, pruning, and knowledge distillation.
- Deploy models to edge devices, mobile environments, and local servers while meeting strict latency targets.
- Build end-to-end MLOps pipelines from data ingestion through deployment.
- Monitor model accuracy, latency, and CPU/GPU utilization in production.
- Evaluate model quality with benchmarking frameworks and custom evaluation suites.
Requirements
- Experience training and fine-tuning small language models with Hugging Face and adapter methods.
- Knowledge of quantization, pruning, knowledge distillation, and lightweight model optimization.
- Experience deploying models to edge devices, mobile environments, or local servers.
- Ability to build end-to-end MLOps pipelines.
- Ability to track model accuracy, latency, and CPU/GPU usage in production.
- Preferred qualifications include edge or mobile deployment experience, ONNX export and cross-platform inference knowledge, and familiarity with experiment tracking, model registries, and CI/CD tooling for ML.
Categories
About Egnyte
Egnyte builds a secure content collaboration and governance platform for mid-market and enterprise teams, delivered as a SaaS subscription. Its products cover file sharing, data classification and protection, compliance workflows, and ransomware detection across cloud and on-premises repositories, with solutions for AEC, financial services, and life sciences. Founded in 2008 and headquartered in Mountain View, it is privately held under GI Partners ownership and serves 22,000+ organizations.
