
Site Reliability Engineer (SRE) - Engineering Productivity
Arista Networks15 hours ago
Responsibilities
- Build, deploy, and operate critical production systems with emphasis on scalability, reliability, observability, performance, and security.
- Monitor, support, and enhance developer experience across internal services and engineering productivity systems.
- Create automation to reduce operational toil and improve production operations.
- Configure monitoring, alerts, and automated alert handling; create and maintain incident response runbooks.
- Design and deploy new systems in staged, scalable, fault-tolerant, and observable ways.
- Triage infrastructure and platform issues, support software engineering triage, and coordinate with third-party vendors.
- Write postmortems and implement solutions to prevent recurring incidents.
- Plan and communicate production maintenance windows.
- Partner with product development teams to identify infrastructure bottlenecks and implement improvements.
- Survey and adopt infrastructure and platform best practices and study open-source systems to improve triage and issue resolution.
Requirements
- BSc in Computer Science or Engineering plus five years of experience, MS in Computer Science or Engineering plus five years of experience, or equivalent work experience.
- Knowledge of Go, Python, or shell scripting sufficient to implement medium-complexity automation workflows.
- Knowledge of Linux or UNIX administration and debugging.
- Hands-on experience operating software systems, infrastructure, or complex applications at scale.
- Experience with server provisioning, particularly from storage and networking perspectives.
- Strong problem-solving and software troubleshooting skills.
- Experience with infrastructure-as-code.
- Desirable experience managing databases such as MariaDB, PostgreSQL, or MongoDB.
- Desirable experience with Docker and virtualization technologies such as KVM, QEMU, or Kata Containers.
- Desirable experience managing monitoring stacks such as Prometheus, Loki, Tempo, InfluxDB, Grafana, or Thanos.
- Desirable experience managing Elasticsearch clusters, Artifactory, Docker registries, CI/CD systems, version control systems, and storage infrastructure.
- Desirable experience with Argo CD, Spinnaker, Perforce, Gerrit, Ansible, NAS, SAN, Ceph, and large Java applications.
Benefits
- Hybrid cloud and globally distributed engineering environment.
- Flat, streamlined management structure with substantial project ownership.
- Opportunities to work across domains and collaborate with product development teams.
- Access to engineering offices and teams across the United States, Australia, Canada, India, and Ireland.
- Inclusive, engineering-centric culture focused on invention, quality, respect, and fun.
Tech Stack
AnsibleArgo CDDockerElasticsearchGoGoogle CloudGrafanaInfluxDBJavaJenkinsKubernetesLinuxMariaDBMongoDBMySQLPostgreSQLPrometheusPythonSpinnaker
Categories
DevOpsSite Reliability
About Arista Networks
Arista Networks is an industry leader in data-driven, client to cloud networking for large data center/AI, campus and routing environments. Arista’s award-winning platforms deliver availability, agility, automation, analytics and security through an advanced network operating stack. visit https://www.arista.com. Additional information and resources can be found at: www.arista.com www.twitter.com/aristanetworks www.facebook.com/AristaNW www.youtube.com/user/AristaNetworks