Responsibilities
- Drive service reliability, availability, and performance across multi-cloud environments by establishing SLOs, SLIs, error budgets, and reliability standards.
- Design, build, and operate enterprise ML platform infrastructure using Dataiku, Amazon SageMaker AI, Databricks, and Google Vertex AI.
- Develop AI-powered observability capabilities using anomaly detection, predictive analytics, and automated remediation.
- Implement, monitor, and optimize LLM, SLM, RAG, and AI agent platforms for performance, governance, scalability, and operational excellence.
- Design Infrastructure as Code, CI/CD pipelines, self-healing systems, and platform automation solutions.
- Architect enterprise ChatOps solutions integrating operational events, observability platforms, AI workflows, and automated remediation.
- Partner with Data Science, AI Engineering, and Platform teams to deliver secure, scalable, production-ready AI/ML solutions.
- Evaluate emerging AI-native operational technologies and recommend improvements to platform reliability and engineering efficiency.
- Assess technical debt and architectural risks and provide strategic recommendations for platform modernization.
- Serve as a technical leader, trusted advisor, mentor, and influencer across reliability engineering, MLOps, cloud platform strategy, and AI-enabled operations.
Requirements
- Bachelor’s degree in Computer Science, Information Technology, Engineering, Data Science, Artificial Intelligence, or a related field is required; a master’s degree is preferred.
- 6-8 years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or related technology areas with enterprise-scale delivery experience.
- Hands-on experience operating across at least two major cloud platforms, including AWS, GCP, and Azure.
- Deep expertise with Databricks, Amazon SageMaker AI, Dataiku, and Google Vertex AI.
- Experience implementing end-to-end ML workflows including model training, experiment tracking, deployment, monitoring, and pipeline orchestration.
- Experience applying anomaly detection, predictive analytics, time-series modelling, and operational intelligence in enterprise environments.
- Advanced proficiency with Terraform, Pulumi, AWS CDK, and modern CI/CD automation practices.
- Strong programming and scripting skills in Python, Go, Bash, or similar languages.
- Experience designing enterprise observability solutions using Prometheus, Grafana, Datadog, and OpenTelemetry, including distributed tracing, logging, and monitoring.
- Experience with automated remediation, AI-assisted operations, enterprise ChatOps, operational workflow automation, and AI-enabled integrations.
- Ability to identify technical debt, assess platform risks, influence technical strategy, and drive modernization initiatives.
- Experience using AI tools, AI agents, and LLM-powered assistants for engineering operations, incident management, and developer productivity.
- Experience with Kubernetes and container orchestration platforms such as EKS, GKE, or AKS is preferred.
- Familiarity with Kubeflow, Feast, MLflow, LangSmith, RAGAS, Evidently AI, or Weights & Biases is preferred.
- Knowledge of cloud cost optimization, policy-as-code, compliance automation, FinOps practices, and multi-cloud governance is preferred.
Benefits
- The role is based in Hyderabad and follows a hybrid work arrangement.
- Regeneron offers a competitive total rewards package that may include bonuses or incentive plans, equity awards, pension or retirement benefits, health and wellness programs, insurance benefits, paid time off, and family support benefits.
- The company provides reasonable accommodation during recruitment where required and conducts legally compliant background checks.
- Many roles are expected to be performed on-site, with specific expectations to be confirmed with the recruiter and hiring manager.
Tech Stack
Categories
About Regeneron
At Regeneron we believe that when the right idea finds the right team, powerful change is possible. As we work across our expanding global network to invent, develop and commercialize life-transforming medicines for people with serious diseases, we’re establishing new ways to think about science, manufacturing and commercialization. And new ways to think about health. Connect with us so we can learn more about you, and you can learn more about our biopharmaceutical medicines. And join us, as we build a future we believe in. Please visit www.regeneron.com/social-media-terms for information on how to engage with us on social media. An important note about privacy: Regeneron is committed to your privacy and will not ask for sensitive personal information such as social security number, date of birth or bank account details via email or social media.
