Global

Site Reliability Engineer

Global
Apply
7 days ago
Berkeley, CA, USAMid Level

Responsibilities

  • Monitor and triage real-time alerts across compute, storage, network, and facility systems.
  • Build preventive automation and develop monitoring-pipeline tools and integrations from APIs through alerts and actions.
  • Support power, cooling, environmental, and other building management systems on the data center floor.
  • Coordinate maintenance activities across teams and accurately track incidents through resolution.
  • Work the overnight shift from 12 a.m. to 8 a.m., five days per week.

Requirements

  • Demonstrated comfort working overnight, five days per week, in a hybrid onsite role in Berkeley, California.
  • Strong Linux and command-line experience, including SSH.
  • Programming or scripting experience with Python, C, C++, Perl, or Java.
  • Network security fundamentals, including ACLs and firewalls.
  • Ability to work independently, communicate across teams, and resolve complex ambiguous problems.
  • Willingness to learn Kubernetes, Prometheus or VictoriaMetrics, Alertmanager, and building management or cooling systems.
  • Preferred experience deploying agentic AI or autonomous automation for technical workflows.
  • Preferred ServiceNow implementation experience and knowledge of ITSM best practices.

Benefits

  • Hybrid work arrangement in Berkeley, California.
  • One-year contract assignment with possible extension based on performance and organizational needs.
  • Overnight schedule from 12 a.m. to 8 a.m., five days per week.

Tech Stack

Categories

Site Reliability
Global

About Global

1,001-5,000 employees
Contact me