3 months ago
Base Salary
$220k - $325k/yr
Responsibilities
- Architect and improve automation, infrastructure-as-code, CI/CD pipelines, and self-healing systems.
- Optimize Kubernetes, Docker, and GCP deployments for performance, capacity, scalability, and low latency.
- Improve build, test, deployment, local development, and production monitoring systems for engineering teams.
- Build and maintain shared tooling and integrate pipelines with third-party vendors.
- Debug difficult technical problems and improve system robustness, operability, and diagnosability.
- Review designs and own their security, scalability, and operational integrity.
- Lead incident response and improve monitoring, observability, and reliability practices.
- Mentor and educate engineers across experience levels and promote reliability as a core engineering value.
Requirements
- 8–10 years of experience in infrastructure engineering or similar DevOps, systems engineering, or site reliability engineering roles.
- Strong programming skills in Python or Go and a record of writing high-quality, well-tested code.
- Deep understanding of distributed systems and experience designing, building, scaling, and maintaining production services.
- Experience with Kubernetes, cloud-native technologies, infrastructure as code, and configuration management tools.
- Experience implementing monitoring and observability solutions, debugging systems, and tuning performance.
- Strong incident management and incident response leadership experience.
- Strong written, verbal, and interpersonal communication skills, including the ability to work with engineers from junior through principal levels.
- Bonus: deep experience with GCP, Prometheus, Grafana, or Datadog; experience with high-throughput and low-latency systems, Go, Terraform, rapid-growth environments, blog posts, or training materials.
Benefits
- Full-time employee benefits include competitive salary and equity, with compensation figures not specified.
- 401(k) program with a 4% match for US employees.
- Health, dental, vision, life, short-term disability, and long-term disability insurance.
- Paid parental, medical, and caregiver leave.
- Flexible time off and holidays.
- Commuter benefits, monthly wellness stipend, in-office setup reimbursement, office amenities, and quarterly team gatherings.
- Full-time role based in the Foster City, California office with in-office work required Monday, Wednesday, and Friday.
- Autonomous work environment.
Tech Stack
Categories
About Replit
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide and over 500,000 business users, Replit is democratizing software development by removing traditional barriers to application creation.
