about 3 hours ago
Responsibilities
- Design, build, and operate large-scale cloud infrastructure and production services.
- Participate in an on-call rotation supporting highly available customer-facing systems.
- Lead incident response efforts and drive post-incident reviews focused on systemic improvements.
- Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
- Partner with engineering teams to improve service availability, scalability, performance, and resilience.
- Continuously improve observability through metrics, logging, tracing, dashboards, and alerting.
- Develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies.
- Eliminate operational toil through automation, tooling, and platform engineering.
- Improve deployment safety and operational workflows through CI/CD and GitOps practices.
- Lead complex reliability initiatives spanning multiple engineering teams.
- Guide engineers in adopting operational best practices and reliability engineering principles.
- Mentor engineers through technical collaboration, design reviews, incident analysis, and knowledge sharing.
- Drive projects from conception through production rollout and long-term operational ownership.
- Explore and apply AI-assisted engineering techniques to improve operational efficiency.
Requirements
- Strong experience operating large-scale production services in AWS and/or GCP.
- Deep expertise with Kubernetes in production environments.
- Experience troubleshooting Kubernetes networking, storage, scheduling, scaling, and workload lifecycle issues.
- Extensive experience with Infrastructure as Code technologies such as Terraform and Helm.
- Strong software engineering skills in Golang and/or Python.
- Experience operating and troubleshooting distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, or Cassandra.
- Strong understanding of cloud networking fundamentals including DNS, load balancing, ingress, and traffic management.
- Experience with observability platforms, monitoring strategies, and production telemetry.
- Understanding of cloud security fundamentals, IAM, and secure infrastructure design.
- Demonstrated success leading complex engineering initiatives across multiple teams.
Benefits
- Immersive, in-person onboarding experience designed to accelerate your impact.
- Support for well-being and social impact initiatives.
- Opportunities for talent development and fostering community connections.
Tech Stack
Apache CassandraAWSDatadogGitGoGoogle Cloud PlatformHelmKubernetesMySQLPostgreSQLPythonRedisSplunkTerraform