
Staff Platform & Reliability Engineer
Interface AI2 days ago
Base Salary
$240k - $320k/yr
Responsibilities
- Define customer-facing SLIs and SLOs, operate error budgets, and govern release decisions through reliability targets.
- Own regional failover, disaster recovery, RTO/RPO definitions, third-party dependency resilience, and tested recovery procedures.
- Build and operate GitOps-based progressive delivery with automated analysis, one-click rollback, drift detection, and change auditing.
- Own AWS infrastructure-as-code, Kubernetes, service mesh, capacity, cost, and scale across the platform.
- Lead paging, severity policies, incident-command rotations, status-page automation, post-mortems, and customer SLA reporting.
- Build metrics, logging, distributed tracing, dashboards, alerts, and burn-rate alerting as code.
- Develop AI-native operational automation, safe runbooks, self-healing workflows, and golden-path service templates.
- Own cloud and cluster security, including IAM least privilege, secrets management, admission and network policy, image signing, SBOMs, scanning, and threat-detection feeds.
- Produce infrastructure compliance evidence for SOC 2 Type II, NCUA/FFIEC examinations, and GLBA reviews.
- Protect AI systems against prompt injection, tool-based data exfiltration, tenant-isolation failures, endpoint abuse, and unsafe AI-assisted engineering workflows.
Requirements
- Senior-most hands-on individual contributor experience running production systems with real uptime commitments and pager participation.
- Deep production Kubernetes on AWS, including service mesh and GitOps delivery across many services.
- Experience managing infrastructure-as-code at multi-account scale and reshaping large existing estates.
- Experience building SLO and error-budget practices, burn-rate alerting, and multi-region or disaster-recovery capabilities with tested failover.
- Daily-practice security experience with least-privilege IAM, secrets management, admission and network policies, software supply-chain controls, and auditor-facing evidence.
- Daily fluency with frontier AI tools such as Claude Code and Cursor, including judgment about agent operating boundaries.
- Strong programming ability in TypeScript and/or Python plus Bash, with clear technical and regulatory-facing writing.
- BS or BA in Computer Science required; an MS or PhD is preferred.
- Preferred experience includes real-time voice or telephony systems, streaming or analytical data platforms, durable workflow engines, chaos engineering, regulated industries, LLM threat modeling, cloud cost engineering, and LLM spend attribution.
- Must be San Francisco-based, committed to working onsite, and able to participate in on-call coverage.
Benefits
- 100% paid health, dental, and vision care.
- 401(k), financial wellness perks, daily meals, commuter benefit, monthly wellness stipend, and mental health, wellness, and family benefits.
- Claude Enterprise and frontier AI tools for every employee.
- Discretionary PTO and paid parental leave.
- Onsite work at the brand-new 21st-floor San Francisco office at 44 Montgomery.
Tech Stack
Categories
DevOpsSite Reliability
About Interface AI
Interface AI builds generative AI virtual assistants for community banks and credit unions, powering autonomous customer-service interactions across voice, chat, and employee-assisted channels. The company sells an AI platform and pre-trained agents delivered as enterprise software to financial institutions, integrating with voice systems and APIs. Founded in 2019 and headquartered in San Francisco, it is privately held and raised a $30M Series A in 2024.