about 3 hours ago
Remote, United States or San Francisco, CA, USASenior / Mid Level
H1B Sponsor
Base Salary
$140k - $288k/yr
Responsibilities
- Ensure the reliability, availability, and performance of production infrastructure and platform services.
- Operate and scale Kubernetes platforms, including governance and support for multi-tenant workloads.
- Manage GitOps-based deployment workflows using ArgoCD and Helm.
- Drive infrastructure provisioning and change management through Terraform/Terragrunt.
- Build and support CI/CD automation and deployment workflows using GitHub Actions.
- Lead incident response efforts, root cause analysis, and post-incident improvement initiatives.
- Reduce operational toil through scripting, tooling, and process automation.
- Advance observability practices across logs, metrics, traces, dashboards, and alerting.
- Support secure secrets integration, IAM-aware operations, and platform guardrails.
- Collaborate closely with application, security, and platform teams to improve reliability and delivery outcomes.
Requirements
- 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Infrastructure.
- Strong hands-on experience operating AWS in production environments.
- Deep expertise in Kubernetes, including cluster operations and troubleshooting.
- Proven experience with Kubernetes multi-tenancy and related governance.
- Experience implementing and operating ArgoCD within a GitOps delivery model.
- Strong hands-on experience with Helm.
- Experience with Terraform/Terragrunt for infrastructure provisioning.
- Solid scripting and automation skills using Bash and/or Python.
- Experience building and maintaining CI/CD pipelines, ideally using GitHub Actions.
- Strong troubleshooting skills across Linux, containers, IAM, networking, and distributed systems.
- Experience with monitoring, alerting, and observability in production environments.
- Demonstrated ownership mindset with experience handling incidents and resolving production issues.
- Strong collaboration and communication skills.
- Bachelor’s degree in computer science, engineering, or a related field.
- Demonstrated ability to use AI to improve workflow efficiency.
- Strong track record of critical evaluation and verification of AI-assisted work.
- High integrity and ownership regarding sensitive data and accountability.