6 hours ago
Base Salary
$222k - $326k/yr
Responsibilities
- Re-engineer the core Teleport product to scale globally and optimize routing latency.
- Rewrite portions of the core product to support the cloud product.
- Build and operate monitoring and observability systems that detect production issues and reduce false positives.
- Automate high-toil operational activities and develop supporting tooling.
- Handle patching, scaling, backup and restore, disaster recovery, and related operational challenges.
- Investigate customer outages and incidents.
- Participate in a 24/7/365 on-call rotation.
- Operate and support the observability platform to maintain visibility and reliability.
Requirements
- 5+ years of progressive experience in software engineering and/or SRE/DevOps roles.
- Strong experience with Linux systems, networking, containers, and troubleshooting.
- Solid Go and Kubernetes development experience.
- Experience developing scripts, automation, lightweight programs, product patches, or tooling incorporating AI agents into operational workflows.
- AWS cloud experience is preferred; GCP experience is acceptable.
- Experience with observability tools such as Prometheus, Grafana, and Loki.
- Ability to complete a collaborative Go coding challenge as part of the interview process.
- Experience working in environments where security, correctness, and system invariants are critical.
- Strong communication skills, intellectual curiosity, and a collaborative, low-ego working style.
Benefits
- Extensive health coverage.
- Annual expense budget.
- Rest and recovery policies.
- Retirement savings plans.
- Professional development opportunities.
- Remote-first work arrangement with a required in-person onboarding week in Oakland, California.
- Take-home coding challenge that usually takes about two weeks as part of the hiring process.
Tech Stack
Categories
Site Reliability
About Teleport
Teleport builds an infrastructure access platform that assigns cryptographic identities to people, machines, and workloads to securely reach servers, Kubernetes, databases, applications, and clouds. It offers enterprise subscriptions and a managed cloud service, with roots in an open-source project. Founded in 2015 and headquartered in Oakland, the privately held company focuses on zero-trust access, privileged access management, and policy governance for engineering teams operating multi-cloud and on-prem environments.
