4 months ago
Bengaluru, IndiaMid Level
Responsibilities
- Operate production infrastructure across multiple GKE clusters with high availability, autoscaling, and observability.
- Manage Argo CD GitOps workflows and implement repeatable infrastructure automation with Terraform and Helm.
- Maintain monitoring, logging, alerting, and product-specific service-level objectives for video and AI workloads.
- Support real-time video infrastructure, including media server scaling, WebRTC, low-latency networking, and high-throughput data paths.
- Support AI agent infrastructure, including LLM inference, asynchronous task queues, and healthcare-system integrations.
- Lead or support incident response, cluster upgrades, and disaster recovery procedures.
- Own infrastructure security through least-privilege access controls, secrets hygiene, security hardening, vulnerability scanning, dependency audits, and policy enforcement.
- Implement compliance-aligned controls such as encryption, audit logging, and network segmentation for healthcare data environments.
- Collaborate with product and engineering teams and contribute to infrastructure and platform engineering practices.
Requirements
- Computer Science or Engineering degree, or equivalent practical experience.
- At least 3 years of hands-on production Kubernetes experience.
- Strong knowledge of CI/CD pipelines and GitOps workflows using Argo CD or similar tools.
- Proficiency with Terraform and Helm for infrastructure automation.
- Experience managing open-source monitoring and logging stacks such as Prometheus, Loki, Grafana, and Alertmanager.
- Working knowledge of cloud security principles, including IAM, network policies, pod security, RBAC, and secrets management.
- Comfort with Linux systems, shell scripting, and basic networking, including UDP/TCP behavior relevant to real-time media or distributed systems.
- Preferred experience with large-scale, multi-tenant, or mixed-workload infrastructure; WebRTC, SFUs, TURN/STUN, or media server orchestration; Vault or Sealed Secrets; Trivy, OPA/Gatekeeper, or Falco; GCP and GKE; healthcare compliance or HIPAA; AI/ML inference workloads; asynchronous pipeline infrastructure; and open-source contributions.
- Fluent written and spoken English and a willingness to share infrastructure, security, and platform engineering knowledge.
Benefits
- Work on varied infrastructure spanning real-time video at scale and AI-driven healthcare automation.
- Join a small, high-ownership, engineering-focused startup with opportunities to grow as an individual contributor or team leader.
- Collaborate with engineers experienced in distributed systems, real-time media, AI infrastructure, and platform engineering.
- Work in the office at least three days per week, with an emphasis on in-office collaboration.
