10 hours ago
Base Salary
$140k - $201k/yr
Responsibilities
- Build and maintain automated reliability tooling, infrastructure-as-code, and observability systems.
- Develop monitoring, logging, and alerting frameworks using tools such as Prometheus, Grafana, and OpenTelemetry.
- Implement automated architectural reviews and reliability guardrails for agent-developed applications.
- Design and implement scalable, fault-tolerant systems that meet defined SLIs and SLOs.
- Automate operational tasks and develop self-healing and auto-remediation mechanisms.
- Participate in on-call rotations, lead incident response, conduct post-incident reviews, and drive systemic improvements.
- Improve deployment and release processes through CI/CD pipelines and progressive delivery.
- Lead observability, reliability, and operational readiness reviews.
- Collaborate with Security and Compliance teams on FedRAMP, NIST, and internal policy requirements.
- Create documentation, runbooks, and internal tooling to improve operational maturity.
Requirements
- Bachelor’s degree in Computer Science, Software Engineering, or a related technical field.
- 3–5 years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering.
- 2+ years of hands-on experience managing and scaling services in cloud environments such as AWS, GCP, or Azure.
- 1+ year of proficiency in at least one modern programming language such as Java, Go, Python, Ruby, or JavaScript.
- Strong understanding of Docker and Kubernetes.
- Experience implementing and maintaining CI/CD pipelines and automation frameworks.
- Working knowledge of metrics, tracing, logging, and alerting observability systems.
- Experience building automated recovery, failover, or chaos-engineering systems.
- Familiarity with event-driven architecture, asynchronous processing, distributed systems, load balancing, and performance optimization.
- Exposure to Terraform, Pulumi, Ansible, and GitOps practices.
- Understanding of FedRAMP, SOC2, or NIST 800-53 security and compliance frameworks.
- Strong analytical, troubleshooting, communication, and documentation skills.
- Experience using AI agentic coding assistants and deploying custom AI agents or automated workflows into production environments.
Benefits
- Comprehensive medical, dental, vision, health savings account, flexible spending accounts, commuter benefits, life and AD&D insurance, disability insurance, accident and critical illness insurance, and a 401(k) with company match.
- Parental leave, unlimited paid time off subject to policy terms, and eight company-wide holidays.
- Referral bonus policy, employee assistance program, pet insurance, travel assistance, wellbeing and childcare discounts, benefit advocates, and learning and development benefits.
- Full-time, in-office attendance five days per week in Mountain View, CA or McLean, VA.
Tech Stack
AnsibleAWSAzureDockerGoGoogle Cloud PlatformGrafanaJavaJavaScriptKubernetesPrometheusPythonRubyTerraform
Categories
DevOpsSite Reliability
About ID.me
ID.me builds a secure digital identity network and wallet used for identity proofing, sign-in, and community verification across government, healthcare, and commerce. The company sells authentication and verification services to agencies and enterprises, including NIST 800-63-3 IAL2/AAL2–approved credentials, plus group eligibility checks for discounts. Founded in 2010 and headquartered in McLean, VA, ID.me supports access at 20 U.S. federal and 45 state agencies and 70+ healthcare organizations.
