Arcesium LLC

Principal Engineer - SRE

Arcesium LLC
Apply
2 months ago
Bengaluru, India or Hyderābād, IndiaStaff+

Responsibilities

  • Own the reliability and availability of platform infrastructure spanning Kubernetes clusters, ArgoCD deployments, PostgreSQL databases, and AWS services.
  • Plan and execute infrastructure maintenance, upgrades, patching, and validation with minimal client impact.
  • Improve CI/CD pipelines and developer tooling, including automated and auditable Terraform workflows.
  • Build and enhance monitoring, alerting, and observability to reduce detection and resolution times.
  • Evaluate infrastructure tooling and capabilities for service mesh, cost optimization, capacity planning, reliability, and scalability.
  • Support convergence and standardization of Limina infrastructure with Arcesium’s broader infrastructure ecosystem.
  • Establish and participate in global on-call rotations and sustainable operational coverage.
  • Troubleshoot complex issues across application, infrastructure, and cloud layers.
  • Write and review infrastructure-as-code, automation scripts, internal tooling, design documents, RFCs, runbooks, and operational procedures.
  • Provide hands-on technical leadership for reliability architecture and large-scale production systems.

Requirements

  • Bachelor’s or master’s degree in computer science, engineering, or a related field.
  • 9+ years of professional engineering experience, including significant principal-level or equivalent individual contributor experience.
  • Hands-on experience in site reliability, infrastructure, DevOps, or platform engineering roles.
  • Strong experience operating production Kubernetes environments, including cluster lifecycle, networking, storage, RBAC, and troubleshooting.
  • Experience running critical workloads on AWS, including EC2, EKS, RDS, IAM, VPC, S3, and CloudWatch.
  • Experience with Datadog, Prometheus, Grafana, ELK, Terraform, CloudFormation, GitLab CI, and Jenkins.
  • Strong Python programming and scripting skills; Bash, Go, or Java familiarity is a plus.
  • Experience with GitOps deployment workflows such as ArgoCD or FluxCD.
  • Production experience with PostgreSQL or similar relational databases, including backup, recovery, upgrades, and performance tuning.
  • Experience with on-call rotations, incident management, post-incident reviews, and large-scale reliability initiatives such as HA/DR, observability, and automation platforms.

Tech Stack

AWSBashDatadogGitLab CI/CDGoGrafanaJavaJenkinsKubernetesPostgreSQLPrometheusPythonTerraform

Categories

DevOpsSite Reliability
Arcesium LLC

About Arcesium LLC

1,001-5,000 employees

Arcesium builds cloud-native software and data platforms for investment managers, hedge funds, private markets, and sell-side firms, spanning front-to-back operations, data warehousing, reconciliation, treasury, and middle-/back-office services. The company operates a SaaS and managed-services model. Founded in 2015 and headquartered in New York City, Arcesium originated from technology developed at the D. E. Shaw group and launched with Blackstone; J.P. Morgan later made a strategic investment.

Contact me