Principal Engineer - SRE
Arcesium LLCabout 2 hours ago
Responsibilities
- Own the reliability and availability of Kubernetes clusters, ArgoCD deployments, PostgreSQL databases, and supporting AWS services.
- Plan and execute infrastructure maintenance, upgrades, patching, and validation with minimal client impact.
- Improve CI/CD pipelines and developer tooling, including automated and auditable Terraform workflows.
- Build and enhance infrastructure monitoring, alerting, and observability to reduce detection and resolution times.
- Evaluate infrastructure tooling and capabilities such as service mesh, cost optimization, and capacity planning.
- Support infrastructure convergence and standardization across tooling, networking, security, and operational practices.
- Establish and participate in on-call rotations across global time zones.
- Troubleshoot complex issues across application, infrastructure, and cloud layers.
- Write and review infrastructure-as-code, automation scripts, internal tooling, design documents, RFCs, runbooks, and operational procedures.
- Lead by example in architecture, code quality, operational maturity, and institutional knowledge sharing.
Requirements
- A bachelor’s or master’s degree in computer science, engineering, or a related field.
- 9+ years of professional engineering experience, including significant time in a principal-level or equivalent individual contributor role.
- Hands-on experience in site reliability engineering, infrastructure engineering, DevOps, or platform engineering.
- Strong experience owning and operating production Kubernetes environments, including cluster lifecycle, networking, storage, RBAC, and troubleshooting.
- Hands-on experience running critical workloads on AWS, including EC2, EKS, RDS, IAM, VPC, S3, and CloudWatch.
- Experience with Datadog, Prometheus, Grafana, ELK, modern alerting design, Terraform, CloudFormation, GitLab CI, and Jenkins.
- Strong programming and scripting ability in Python; Bash, Go, or Java familiarity is a plus.
- Experience with GitOps-based deployment workflows such as ArgoCD or FluxCD.
- Production experience with PostgreSQL or similar relational databases, including backup, recovery, upgrades, and performance tuning.
- Experience with on-call rotations, incident management, and post-incident review processes.
- A proven track record designing and delivering large-scale reliability initiatives involving high availability, disaster recovery, observability, or automation platforms.