6 hours ago
Base Salary
$222k - $326k/yr
Responsibilities
- Re-engineer the core Teleport product to scale globally and optimize routing latency.
- Rewrite portions of the product to support Teleport Cloud.
- Build monitoring and observability systems that detect production issues while minimizing false positives.
- Automate high-toil operational activities and develop tooling, including tooling that incorporates AI agents into operational workflows.
- Handle patching, scaling, backup and restore, disaster recovery, and related operations.
- Investigate customer-impacting outages and incidents.
- Participate in the 24/7/365 on-call rotation.
- Operate and support the observability platform to maintain visibility and reliability.
Requirements
- 8+ years of progressive experience in software engineering and/or SRE/DevOps roles.
- Strong experience with Linux systems, networking, containers, and troubleshooting.
- Solid Go and production Kubernetes development experience.
- Technical leadership experience.
- Experience developing scripts and automation, submitting patches to product codebases, or building operational tooling incorporating AI agents.
- AWS cloud experience preferred; GCP experience acceptable.
- Experience with observability tools such as Prometheus, Grafana, and Loki.
- Ability to reason about correctness and system invariants, including formal or property-based methods.
- Willingness to complete a collaborative Go coding challenge during the interview process.
- Strong communication, intellectual curiosity, transparency, honesty, and a no-ego mindset.
Benefits
- Remote-first work arrangement with an in-person onboarding week in Oakland, California.
- Extensive health coverage.
- Annual expense budget.
- Rest and recovery policies.
- Retirement savings plans.
- Professional development opportunities.
- The take-home coding challenge usually takes about two weeks.
Tech Stack
Categories
Site Reliability
About Teleport
Teleport builds an infrastructure access platform that assigns cryptographic identities to people, machines, and workloads to securely reach servers, Kubernetes, databases, applications, and clouds. It offers enterprise subscriptions and a managed cloud service, with roots in an open-source project. Founded in 2015 and headquartered in Oakland, the privately held company focuses on zero-trust access, privileged access management, and policy governance for engineering teams operating multi-cloud and on-prem environments.
