2 months ago
Remote, United StatesSenior
Base Salary
$222k - $342k/yr
Responsibilities
- Re-engineer Teleport’s core product to scale globally and optimize routing latency.
- Rewrite portions of the core product to support the cloud offering.
- Build monitoring and observability capabilities that identify production issues while minimizing false positives.
- Automate high-toil operational activities.
- Handle operational work including patching, scaling, backup and restore, and disaster recovery.
- Investigate customer outages and incidents.
- Operate and support the observability platform to maintain visibility and reliability.
- Participate in a 24/7/365 on-call rotation.
Requirements
- Strong experience with Linux systems, networking, containers, and troubleshooting.
- Solid Go and Kubernetes development experience.
- Experience developing scripts, automation, lightweight programs, product patches, or operational tooling incorporating AI agents.
- Preferred AWS cloud experience; GCP experience is acceptable.
- Experience with observability tools such as Prometheus, Grafana, and Loki.
- Experience operating and supporting observability platforms.
- Experience working in environments where security, correctness, and system invariants are critical.
- Willingness to complete and collaboratively discuss a Go coding challenge during the interview process.
- Excellent communication skills, intellectual curiosity, transparency, honesty, and a collaborative no-ego mindset.
Benefits
- Remote-first and globally distributed work arrangement
- In-person onboarding required in the Oakland, California office for one week
- Extensive health coverage
- Annual expense budget
- Rest and recovery policies
- Retirement savings plans
- Professional development opportunities
- Background checks are conducted
- Interview coding challenge typically takes about two weeks
Tech Stack
Categories
DevOpsSite Reliability
About Teleport
Teleport is the AI Infrastructure Identity Company, modernizing identity, access, and policy for infrastructure, improving engineering velocity and infrastructure resiliency against human factors and compromise.
