4 days ago
Santa Clara, CA, USASenior / Staff+
Base Salary
$155k - $230k/yr
Responsibilities
- Design, build, and operate scalable infrastructure across cloud, Kubernetes, on-premises, and hybrid environments.
- Define and implement solutions for complex infrastructure and engineering productivity problems from architecture through production.
- Build internal developer platforms, tools, and services that improve development, deployment, and operational workflows.
- Design and optimize CI/CD platforms and pipelines, including build and test performance, reliability, scalability, and developer feedback.
- Automate provisioning, configuration, deployment, testing, monitoring, and operational workflows.
- Develop reusable Infrastructure as Code, automation, and platform capabilities.
- Build reliable, observable, resilient, secure, and disaster-recovery-capable systems.
- Troubleshoot infrastructure and platform issues, participate in incident response and root-cause analysis, and resolve systemic problems.
- Evaluate technologies, recommend architectures, establish standards, and lead technical initiatives across multiple engineering teams.
Requirements
- 6+ years of experience in infrastructure, platform engineering, software engineering, SRE, DevOps, or related fields.
- Strong programming and software engineering fundamentals with experience developing production-quality tools and services.
- Deep production Kubernetes experience, including troubleshooting complex environments.
- Experience with AWS, Azure, or GCP and/or on-premises infrastructure, including self-managed or bare-metal environments.
- Experience with Infrastructure as Code such as Terraform or Ansible and with CI/CD platforms and deployment automation.
- Strong understanding of Linux, networking, distributed systems, infrastructure architecture, server provisioning, configuration management, patching, upgrades, and lifecycle management.
- Hands-on experience with Ansible, Chef, or similar configuration management and automation technologies.
- Strong programming skills in at least one of Go, Python, Rust, or C++.
- Experience with observability, monitoring, logging, and reliability engineering practices.
- Ability to lead complex technical initiatives, evaluate tradeoffs, and drive scalable solutions through production.
- Strong communication and collaboration skills across technical and cross-functional teams.
- Preferred experience includes internal developer platforms, Kubernetes operators or controllers, Kubernetes CNI/CSI, distributed technologies, Jenkins, VPN, SSO, infrastructure security, network infrastructure, data center operations, and enterprise customer deployments.
Benefits
- Hybrid work arrangement in Santa Clara, California.
- Competitive base salary, equity, medical, dental, vision, life insurance, retirement savings, wellness program, disability coverage, holidays, and other benefits.
- Unlimited PTO and 40 hours of Volunteer Time Off per year.
- Internet stipend and 401(k).
Tech Stack
AnsibleApache CassandraApache KafkaAWSAzureC++ChefElasticsearchGoGoogle Cloud PlatformJenkinsKubernetesLinuxPythonRustTerraform
