GrepJob
Arcesium LLC

Principal Engineer - SRE

Arcesium LLC
Apply
about 2 hours ago
Bengaluru, India or Hyderābād, IndiaStaff+
H1B Sponsor

Responsibilities

  • Own the reliability and availability of Kubernetes clusters, ArgoCD deployments, PostgreSQL databases, and supporting AWS services.
  • Plan and execute infrastructure maintenance, upgrades, patching, and validation with minimal client impact.
  • Improve CI/CD pipelines and developer tooling, including automated and auditable Terraform workflows.
  • Build and enhance infrastructure monitoring, alerting, and observability to reduce detection and resolution times.
  • Evaluate infrastructure tooling and capabilities such as service mesh, cost optimization, and capacity planning.
  • Support infrastructure convergence and standardization across tooling, networking, security, and operational practices.
  • Establish and participate in on-call rotations across global time zones.
  • Troubleshoot complex issues across application, infrastructure, and cloud layers.
  • Write and review infrastructure-as-code, automation scripts, internal tooling, design documents, RFCs, runbooks, and operational procedures.
  • Lead by example in architecture, code quality, operational maturity, and institutional knowledge sharing.

Requirements

  • A bachelor’s or master’s degree in computer science, engineering, or a related field.
  • 9+ years of professional engineering experience, including significant time in a principal-level or equivalent individual contributor role.
  • Hands-on experience in site reliability engineering, infrastructure engineering, DevOps, or platform engineering.
  • Strong experience owning and operating production Kubernetes environments, including cluster lifecycle, networking, storage, RBAC, and troubleshooting.
  • Hands-on experience running critical workloads on AWS, including EC2, EKS, RDS, IAM, VPC, S3, and CloudWatch.
  • Experience with Datadog, Prometheus, Grafana, ELK, modern alerting design, Terraform, CloudFormation, GitLab CI, and Jenkins.
  • Strong programming and scripting ability in Python; Bash, Go, or Java familiarity is a plus.
  • Experience with GitOps-based deployment workflows such as ArgoCD or FluxCD.
  • Production experience with PostgreSQL or similar relational databases, including backup, recovery, upgrades, and performance tuning.
  • Experience with on-call rotations, incident management, and post-incident review processes.
  • A proven track record designing and delivering large-scale reliability initiatives involving high availability, disaster recovery, observability, or automation platforms.

Tech Stack

AWSBashDatadogGitLab CI/CDGoGrafanaJavaJenkinsKubernetesPostgreSQLPrometheusPythonTerraform

Categories