2 days ago
Bethesda, MD, USA or Richardson, TX, USASenior
Base Salary
$100k - $215k/yr
Responsibilities
- Design, develop, and operate automation, self-service tools, dashboards, data pipelines, APIs, shared services, and integrations for incident management and operational reliability.
- Lead technical response during high-severity incidents, including troubleshooting strategy, impact analysis, cross-team coordination, and safe service restoration.
- Conduct post-incident reviews, root cause analysis, corrective action planning, and systemic reliability improvements.
- Create and maintain runbooks, readiness criteria, triage models, resilience practices, and operational controls.
- Apply engineering standards for deployment, testing, observability, security, production readiness, infrastructure as code, rollback, and operational support.
- Contribute to architecture and design reviews and partner with SRE, platform, product, infrastructure, security, and business stakeholders.
- Mentor engineers through technical leadership, code and design reviews, documentation, and operational coaching.
- Participate in a 24x7 on-call rotation for mission-critical platforms and incident response.
Requirements
- At least 4 years of professional software engineering experience, preferably in platform engineering, reliability engineering, backend engineering, distributed systems, or operational tooling.
- At least 3 years of experience with architecture, design, system reliability, scalability, and technical delivery for production systems.
- At least 2 years of experience with open-source frameworks, modern engineering practices, or platform technologies.
- At least 2 years of experience with Azure, AWS, GCP, another cloud provider, or equivalent complex hybrid environments.
- Hands-on proficiency in multiple languages, including Go, Java, Python, or C#, for production-grade applications on Kubernetes and serverless technologies such as KNative.
- Experience with SQL and NoSQL technologies, cloud-native data services, data pipelines, analytics, dashboards, observability, and incident management platforms.
- Experience with OpenTelemetry and observability platforms such as Grafana, Datadog, Splunk, or Azure Monitor.
- Experience improving incident management, post-incident review, or reliability processes through automation, data, and cross-team collaboration.
- Strong incident forensics, root cause analysis, observability, reliability engineering, system design, and distributed-systems skills.
- Experience supporting high-severity production incidents and owning production systems in 24x7 environments.
- Proficiency with AI-assisted development tools such as Claude Code, Cursor, and GitHub Copilot.
- Bachelor’s degree in Computer Science, Information Systems, or equivalent education or work experience.
- Strong communication, technical leadership, stakeholder influence, and presentation skills, including the ability to work effectively under pressure.
Benefits
- Competitive pay, benefits, and flexibility to support employee well-being and future needs.
- Personalized development programs, mentorship, and certification assistance.
- Inclusive and collaborative culture rooted in shared success.
- Participation in a 24x7 on-call rotation and production support for mission-critical platforms is required.
- GEICO states that it will not sponsor a new applicant for employment authorization for this position.
Tech Stack
Apache SparkApache SupersetAWSAzureC#DatadogGoGoogle Cloud PlatformGrafanaJavaKubernetesPythonSplunkSQL
Categories
BackendSite Reliability
About GEICO
GEICO sells auto and other personal lines insurance to U.S. consumers through a primarily direct-to-consumer model via web, mobile, and phone, with some local agents. Its products include car, motorcycle, RV, boat, homeowners, renters, condo, umbrella, and commercial auto coverage, plus roadside assistance. Founded in 1936 and headquartered in Bethesda, Maryland, GEICO is a subsidiary of Berkshire Hathaway and is among the largest auto insurers in the United States.
