2 days ago
Bethesda, MD, USA or Richardson, TX, USAStaff+
Base Salary
$120k - $260k/yr
Responsibilities
- Design, develop, and operate automation, self-service tools, dashboards, and data pipelines for incident management, on-call, paging, and troubleshooting.
- Build shared services, APIs, data contracts, automation, and integrations that standardize incident response and reduce operational risk.
- Set engineering standards for design, deployment, testing, observability, security, operational support, and production readiness.
- Lead technical response during high-severity incidents, including troubleshooting strategy, cross-team coordination, impact analysis, and service restoration.
- Lead or significantly influence post-incident reviews, root cause analysis, corrective action planning, runbooks, readiness criteria, and resilience practices.
- Lead architecture and design reviews across teams, mentor engineers, and influence technical direction and reliability culture.
Requirements
- 10+ years of professional software engineering experience, preferably in platform, reliability, backend, distributed-systems, or operational-tooling environments.
- 8+ years of experience with architecture, design, system reliability, scalability, and technical leadership for production systems.
- 6+ years of experience with open-source frameworks, modern engineering practices, or platform technologies.
- 4+ years of experience with Azure, AWS, GCP, another cloud provider, or equivalent complex hybrid-environment experience.
- Hands-on proficiency with Go, Java, Python, and C# for production-grade applications on Kubernetes and Knative in Azure and AWS.
- Experience with SQL and NoSQL technologies, cloud-native data services, data pipelines, analytics, dashboards, OpenTelemetry, observability platforms, and PagerDuty.
- Experience with Spark, Trino, Grafana, Superset, Power BI, Datadog, Splunk, Azure Monitor, Claude Code, Cursor, and GitHub Copilot.
- Demonstrated experience improving incident-management, post-incident review, or reliability processes at scale through automation, data, and cross-team influence.
- Strong incident forensics, root cause analysis, observability, reliability engineering, system design, distributed-systems, communication, and technical coaching skills.
- Bachelor's degree in Computer Science, Information Systems, or equivalent education or work experience.
- Ability to participate in a 24x7 on-call rotation and perform effectively during high-severity incidents and other stressful production-support situations.
Benefits
- Annual salary range of $120,000.00-$260,000.00.
- Personalized development programs, mentorship, and certification assistance.
- Inclusive and collaborative culture with shared-success values.
- Benefits and flexibility intended to support employee well-being and future needs.
- Participation in a 24x7 on-call rotation is required.
Tech Stack
Apache SparkApache SupersetAWSAzureC#DatadogGoGoogle Cloud PlatformGrafanaJavaKubernetesPythonSplunkSQL
Categories
Site Reliability
About GEICO
GEICO sells auto and other personal lines insurance to U.S. consumers through a primarily direct-to-consumer model via web, mobile, and phone, with some local agents. Its products include car, motorcycle, RV, boat, homeowners, renters, condo, umbrella, and commercial auto coverage, plus roadside assistance. Founded in 1936 and headquartered in Bethesda, Maryland, GEICO is a subsidiary of Berkshire Hathaway and is among the largest auto insurers in the United States.
