1 day ago
Bengaluru, IndiaSenior / Staff+
Responsibilities
- Own and operate scalable, reliable, and cost-efficient AWS infrastructure for the Network Assurance Data Platform.
- Manage and optimize data workflows using Apache Airflow, AWS EMR, Spark, and Hadoop-based platforms.
- Operate Amazon EKS and containerized environments supporting data and ML workloads.
- Build infrastructure automation and operational tooling with Terraform and Python.
- Drive FinOps practices including cost visibility, allocation, forecasting, anomaly detection, optimization, and governance.
- Improve observability, alerting, incident response, capacity planning, autoscaling, storage optimization, and operational readiness.
- Identify infrastructure, performance, workload, and cost issues and convert analysis into technical recommendations.
- Partner with data engineering, ML, platform, finance, product, and engineering teams to improve reliability, scalability, performance, and cost efficiency.
- Lead technical discussions, influence architecture decisions, mentor engineers, and drive continuous operational improvement.
Requirements
- Bachelor’s degree or higher in Engineering, Computer Science, or equivalent practical experience.
- 8–10 years of relevant experience in Site Reliability Engineering, DevOps, cloud infrastructure, platform engineering, data infrastructure, or production engineering.
- Strong hands-on experience operating production infrastructure on AWS, including Apache Airflow, AWS EMR, Spark or Hadoop, Amazon EKS or Kubernetes, Terraform, and Python.
- Experience supporting containerized workloads, ML infrastructure, data platforms, workflow orchestration, and production troubleshooting.
- Practical experience with cloud cost optimization, FinOps, AWS cost analysis, tagging, budgeting, forecasting, and cost governance.
- Strong understanding of Linux systems, networking, distributed systems, observability, incident management, root cause analysis, and reliability engineering.
- Experience with observability platforms such as CloudWatch, Prometheus, Grafana, Splunk, OpenSearch, Datadog, or similar tools.
- Preferred experience includes large-scale SaaS or high-volume data infrastructure, ML infrastructure, batch processing, and optimization of EMR, Spark, Airflow, EKS, storage, and compute workloads.
- Preferred AWS experience includes EC2, S3, RDS, IAM, VPC, CloudWatch, OpenSearch, Lambda, ElastiCache, Savings Plans, Reserved Instances, Spot adoption, Graviton migration, and storage lifecycle management.
- Preferred experience includes AWS Cost Explorer, AWS CUR, Cloudability, CloudHealth, Kubecost, CI/CD systems, GitHub workflows, Atlantis, Puppet, Ansible, Helm, and Argo CD.
- FinOps certification or equivalent hands-on cloud financial management experience is a plus.
- Ability to operate independently, lead initiatives end to end, communicate clearly, and influence teams toward reliable, scalable, and cost-aware designs.
Benefits
- Hybrid work arrangement in Bangalore, India.
- Opportunity to work on large-scale network assurance, data, ML, cloud infrastructure, reliability, and FinOps systems at Cisco ThousandEyes.
- Cross-functional collaboration with data engineering, ML, platform, finance, product, and engineering teams.
- Technical leadership, mentorship, and opportunities to drive platform and operational improvements.
Tech Stack
AnsibleApache AirflowApache HadoopApache SparkArgo CDAWSDatadogGrafanaHelmKubernetesLinuxPrometheusPuppetPythonSplunkTerraform
Categories
DevOpsSite Reliability
About Cisco
Cisco designs and sells networking, security, and collaboration platforms for enterprises, service providers, and governments, spanning routers and switches, Wi‑Fi, firewalls, zero‑trust, observability, and cloud-managed IT (Meraki) plus Webex. Its business model mixes hardware, software subscriptions, and support/consulting services. Founded in 1984 and headquartered in San Jose, California, Cisco is a public company traded on Nasdaq and serves customers across data centers, campuses, and service provider networks.
