1 day ago
San Jose, CA, USASenior
Base Salary
$105k - $175k/yr
Responsibilities
- Lead and participate in incident response, postmortems, root-cause analysis, gap assessments, escalation, and recovery activities.
- Monitor production environments, respond to alerts, diagnose system and performance issues, and coordinate remediation through completion.
- Design and maintain automation, scripts, integrations, runbooks, workflows, dashboards, and visualizations for infrastructure operations.
- Install, configure, troubleshoot, and support hardware, software, storage, networking, cloud, Kubernetes, and other infrastructure services.
- Establish and improve logging, monitoring, alerting, metrics, tracing, observability, performance analysis, and service-level objective practices.
- Develop recovery procedures and participate in disaster recovery, resilience, business continuity, change management, and operational readiness activities.
- Lead or contribute to Operations Team projects from planning and implementation through documentation, transition to support, and closure.
- Partner with development, security, operations, support teams, vendors, and stakeholders to resolve issues and meet delivery commitments.
- Review and improve technical procedures, scripts, automation, and operational documentation while providing guidance to less-experienced team members.
Requirements
- 5+ years of experience in Site Reliability Engineering, Systems Engineering, DevOps, Infrastructure Engineering, or a related field.
- Bachelor's degree in Engineering, Computer Science, or Information Technology, or equivalent professional experience.
- Demonstrated experience supporting highly available production systems.
- Experience leading incident reviews, postmortems, root-cause analysis, and remediation planning.
- Experience working across infrastructure, application, security, and operations teams.
- Hands-on expertise with cloud and on-premises infrastructure, Kubernetes, containerized workloads, virtualization, and distributed systems.
- Advanced knowledge of Linux/UNIX and Windows environments, storage and file systems, backup, disaster recovery, and business continuity.
- Strong automation and infrastructure engineering skills, including Infrastructure as Code, configuration management, scripting, provisioning, deployments, remediation, and security risk mitigation.
- Strong problem-solving, analytical, organizational, communication, prioritization, and execution skills.
Benefits
- Comprehensive medical, dental, and vision benefits through a multi-carrier program.
- 401(k) with company match and an Employee Share Purchase Plan.
- Wellness platform with incentives, Headspace subscription, Employee Assistance, and time-off programs.
- Short- and long-term disability, life and accidental death insurance, critical illness, and hospital indemnity coverage.
- Family benefits including bonding and family care leaves, adoption, and surrogacy benefits.
- Health Savings, Health Care, Dependent Care, and Commuter Spending Accounts.
- Annual paid time off plus up to two paid days each for Employee Resource Group participation and volunteering.
- Annual incentive bonus eligibility and country-specific benefits.
Tech Stack
Categories
DevOpsSite Reliability
About RELX
RELX provides information-based analytics, research content, and decision tools for scientists, legal and risk professionals, insurers, and governments. Through segments including Elsevier (STM), LexisNexis (legal and risk), and RX (exhibitions), it sells subscriptions, data services, and software platforms used in 180+ countries. Headquartered in London, it is a public company listed in London and Amsterdam.