1 month ago
Remote, IndiaSenior
Responsibilities
- Upgrade ML asset and model management systems to improve developer experience and governance.
- Build and optimize model-serving infrastructure for inference latency and cost efficiency.
- Architect inference pipelines balancing latency, throughput, cost, and hardware acceleration options.
- Implement enterprise-scale, cost-efficient solutions and ML CI/CD pipelines.
- Contribute to architectural decisions for distributed ML systems and evaluate new technologies.
- Collaborate with MLEs, QA Engineers, and DevOps Engineers in a cross-functional distributed team.
Requirements
- 5+ years of software engineering experience with Python.
- Experience with model lifecycle management tools such as MLFlow, Weights & Biases, or equivalent.
- Experience with data management ecosystems covering quality, transformation, and cataloging.
- Experience with ML frameworks, particularly PyTorch.
- Experience optimizing ML models using AWS Neuron, ONNX, or TensorRT.
- Proven experience building and operating AWS serverless architectures, including event-driven processing with SQS/SNS and serverless caching solutions.
- Experience with Docker and orchestration tools.
- Strong knowledge of RESTful API design and implementation.
- Proficiency writing high-quality, secure code and familiarity with static code analysis tools.
- Knowledge of computer science fundamentals, including algorithm design, problem solving, and complexity analysis.
- Preferred experience with model compilation, quantization, performance profiling, and ML inference benchmarking.
- Preferred experience in regulated industries with strict compliance requirements for cloud-native solutions.
- Strong analytical, conceptual, communication, and English language skills.
