Principal Engineer - SRE
Arcesium LLC2 months ago
Bengaluru, India or Hyderābād, IndiaStaff+
Responsibilities
- Own the reliability and availability of platform infrastructure spanning Kubernetes clusters, ArgoCD deployments, PostgreSQL databases, and AWS services.
- Plan and execute infrastructure maintenance, upgrades, patching, and validation with minimal client impact.
- Improve CI/CD pipelines and developer tooling, including automated and auditable Terraform workflows.
- Build and enhance monitoring, alerting, and observability to reduce detection and resolution times.
- Evaluate infrastructure tooling and capabilities for service mesh, cost optimization, capacity planning, reliability, and scalability.
- Support convergence and standardization of Limina infrastructure with Arcesium’s broader infrastructure ecosystem.
- Establish and participate in global on-call rotations and sustainable operational coverage.
- Troubleshoot complex issues across application, infrastructure, and cloud layers.
- Write and review infrastructure-as-code, automation scripts, internal tooling, design documents, RFCs, runbooks, and operational procedures.
- Provide hands-on technical leadership for reliability architecture and large-scale production systems.
Requirements
- Bachelor’s or master’s degree in computer science, engineering, or a related field.
- 9+ years of professional engineering experience, including significant principal-level or equivalent individual contributor experience.
- Hands-on experience in site reliability, infrastructure, DevOps, or platform engineering roles.
- Strong experience operating production Kubernetes environments, including cluster lifecycle, networking, storage, RBAC, and troubleshooting.
- Experience running critical workloads on AWS, including EC2, EKS, RDS, IAM, VPC, S3, and CloudWatch.
- Experience with Datadog, Prometheus, Grafana, ELK, Terraform, CloudFormation, GitLab CI, and Jenkins.
- Strong Python programming and scripting skills; Bash, Go, or Java familiarity is a plus.
- Experience with GitOps deployment workflows such as ArgoCD or FluxCD.
- Production experience with PostgreSQL or similar relational databases, including backup, recovery, upgrades, and performance tuning.
- Experience with on-call rotations, incident management, post-incident reviews, and large-scale reliability initiatives such as HA/DR, observability, and automation platforms.
Tech Stack
Categories
DevOpsSite Reliability
About Arcesium LLC
Arcesium builds cloud-native software and data platforms for investment managers, hedge funds, private markets, and sell-side firms, spanning front-to-back operations, data warehousing, reconciliation, treasury, and middle-/back-office services. The company operates a SaaS and managed-services model. Founded in 2015 and headquartered in New York City, Arcesium originated from technology developed at the D. E. Shaw group and launched with Blackstone; J.P. Morgan later made a strategic investment.