26 days ago
Bengaluru, IndiaStaff+
Responsibilities
- Own the architecture and technical direction for high-throughput real-time data processing and distributed messaging systems.
- Define and execute the ML platform roadmap for model training and serving infrastructure.
- Lead GPU capacity planning, scheduling, utilization, burst handling, and cost management across training and inference fleets.
- Scale model inference through batching, routing, quantization, compilation, autoscaling, and caching while maintaining output quality.
- Drive uptime, observability, and incident response for production serving systems.
- Translate ambiguous business problems into executable technical plans and coordinate cross-functional initiatives.
- Lead design reviews, mentor senior engineers, and establish engineering standards across teams.
- Evaluate emerging tools and techniques and adopt them when their complexity is justified.
Requirements
- 10+ years building backend and infrastructure systems with ownership of architecture and design at scale.
- Deep hands-on experience with large-scale databases, high-throughput messaging systems, and real-time job queues.
- Ability to navigate complex codebases and reason about architectural tradeoffs in systems the candidate did not build.
- Experience mentoring senior engineers and driving technical decisions through influence rather than authority.
- Strong written communication for technical and executive audiences across time zones.
- BTech, MTech, or PhD in Computer Science, or equivalent experience.
- Bonus: production experience with Django, Celery, Redis, PostgreSQL, and Google Cloud.
- Bonus: hands-on experience scaling GPU infrastructure and model inference in production, including capacity planning, scheduling, autoscaling, and latency and cost optimization.
- Bonus: experience with vLLM, TensorRT, Triton, Ray Serve, or equivalent inference tooling.
- Bonus: experience scaling a platform through a comparable growth stage or at a global product company’s India site.
- Bonus: background in speech, NLP, or information retrieval systems.
Benefits
- Health coverage for the employee, spouse, children, and parents.
- Real ownership over systems used by enterprises worldwide and autonomy over how they are built.
- The role is based in Noida or Bengaluru and involves collaboration across India and US-based teams; the hiring process typically takes 2 weeks across 4 stages.
