Base Salary
$180k - $200k/yr
Responsibilities
- Architect and scale observability platforms for real-time health telemetry from thousands of distributed vehicle nodes.
- Develop systems that remain performant across diverse hardware, intermittent connectivity, and rapidly scaling fleets.
- Define alerting and criticality models that distinguish transient anomalies from systemic reliability issues.
- Design detection logic for silent failures such as sensor degradation, compute saturation, and recording pipeline stalls.
- Automate detection, triage, and mitigation to reduce manual intervention.
- Collaborate with Operations and Engineering on safe automated responses to recurring hardware and software failures.
- Build interfaces that help Operations identify issues and Engineering diagnose and deploy mitigations rapidly.
- Lead reliability-focused design reviews and translate operational problems into technical requirements and roadmaps.
- Apply data analytics to identify latent telemetry patterns and proactively detect regressions and hardware degradation.
Requirements
- 5+ years of relevant industry experience in software engineering, site reliability, or systems engineering.
- Experience with observability platforms such as Prometheus, Grafana, or ELK in edge, IoT, or hardware-integrated environments.
- Production coding experience in one or more of Go, Python, or C++.
- Proficiency in Linux internals and shell scripting for debugging edge devices or hardware-adjacent systems.
- Ability to debug across services, Docker containers, and networking stacks.
- A proven track record owning reliability, infrastructure, or platform systems for large-scale production workloads.
- Experience designing and operating metrics, logging, alerting, and dashboard systems.
- Experience defining and implementing SLIs and SLOs for system availability or data yield.
- Deep understanding of networking protocols including TCP/IP, gRPC, or MQTT in bandwidth-constrained environments.
- Experience driving complex technical projects and architectural reviews across multiple teams from design through production.
- Knowledge of sensor data protocols such as Camera, LiDAR, or Radar, or hardware-to-cloud data ingestion pipelines.
- Experience with grey failure detection and management in complex distributed systems.
- Experience with fleet health for large-scale hardware deployments where automation replaced manual intervention.
Benefits
- Base salary range of USD $180,000-$200,000 per year.
- Eligibility for Uber’s bonus program and potential equity awards and other compensation.
- Eligibility for a 401(k) plan and various benefits.
About Uber
We are Uber. The go-getters. The kind of people who are relentless about our mission to help people go anywhere and get anything and earn their way. Movement is what we power. It’s our lifeblood. It runs through our veins. It’s what gets us out of bed each morning. It pushes us to constantly reimagine how we can move better. For you. For all the places you want to go. For all the things you want to get. For all the ways you want to earn. Across the entire world. In real time. At the incredible speed of now. The idea for Uber was born on a snowy night in Paris in 2008, and ever since then our DNA of reimagination and reinvention carries on. We’ve grown into a global platform powering flexible earnings and the movement of people and things in ever expanding ways. We’ve gone from connecting rides on 4 wheels to 2 wheels to 18-wheel freight deliveries. From takeout meals to daily essentials to prescription drugs to just about anything you need at any time and earning your way. From drivers with background checks to real-time verification, safety is a top priority every single day. At Uber, the pursuit of reimagination is never finished, never stops, and is always just beginning.
