2 hours ago
Bengaluru, IndiaSenior
Responsibilities
- Collaborate with peers and stakeholders to understand dependencies, shape solutions, and productionize experimental data science workflows.
- Develop, refactor, test, and optimize complex machine learning and software components using sound software engineering practices.
- Design big data and ML applications, including model training, evaluation, and batch and streaming inference at scale.
- Evaluate, monitor, and operate production ML models using latency, throughput, accuracy, drift, and business KPI metrics.
- Diagnose data and model drift, performance regressions, and pipeline instability, and implement mitigations such as retraining, recalibration, and feature or architecture changes.
- Design production guardrails including safety constraints, thresholds, fallbacks, and safe defaults.
- Improve code, model architecture, memory and compute efficiency, observability, policies, and operational processes.
- Use AI-assisted tools and generative AI or LLM techniques responsibly across the software and ML engineering lifecycle.
- Mentor junior engineers and lead complex, well-defined projects.
Requirements
- 5+ years of relevant professional experience with end-to-end machine learning engineering pipelines in production, including feature engineering, training, validation, deployment, scoring, monitoring, iteration, and streaming applications in hybrid or cloud environments.
- Bachelor’s or Master’s degree in a technical field such as Computer Science, or equivalent relevant work experience.
- Strong command of Spark or similar big data frameworks, including optimization and debugging of large-scale data processing applications.
- Proficiency with PyTorch and/or TensorFlow and experience integrating models into production inference services at scale.
- Strong ML fundamentals, including working knowledge of deep learning and big data concepts, with experience scaling ML models for production latency, throughput, and memory requirements.
- Hands-on experience with production ML model evaluation and monitoring, metric and alert design, feedback loops, and MLOps practices such as experiment tracking, model registries, deployment, and observability tools.
- Familiarity with secure data access and governance, including IAM policies for S3, and with distributed systems for ML training and serving.
- Working knowledge of generative AI and LLM applications such as prompting, RAG, fine-tuning, embeddings, vector stores, and evaluation is strongly preferred.
- Hands-on experience with AI-assisted engineering tools such as GitHub Copilot or Claude Code is preferred.
