8 months ago
Remote, United StatesStaff+
Responsibilities
- Architect the internal developer platform and cloud engineering capabilities across internal and client-facing infrastructure.
- Design and oversee GCP environments, GKE cluster operations, workload management, and public-cloud resources.
- Define and maintain SLOs, SLIs, and error budgets while improving system reliability and availability.
- Lead incident response, on-call rotations, root-cause analysis, and post-mortem reviews.
- Implement and optimize monitoring, alerting, and observability systems using tools such as OpenTelemetry, Prometheus, Grafana, and Honeycomb.
- Drive automation, self-service, security by default, least privilege, resilient systems, and infrastructure engineering practices.
- Mentor platform engineers and enable platform engineering capabilities within software engineering teams.
- Develop and maintain tooling, infrastructure, tests, documentation, runbooks, ADRs, and CI/CD pipelines.
- Collaborate on capacity planning, performance optimization, peer reviews, pair programming, and continuous improvement.
- Promote testing, chaos engineering, load testing, reliability testing, and strong engineering practices across the platform.
Requirements
- Bachelor's degree in Computer Science, Computer Engineering, or a related field, or 8+ years of experience as a software engineer.
- Proficiency with Kubernetes and extensive experience with Unix/Linux; CKA or CKAD certification is optional.
- Proficiency in multiple languages or infrastructure languages, ideally including Golang, Node.js, Python, and HCL.
- Familiarity with at least two of GCP, AWS, and Azure in multi-cloud environments.
- Advanced experience creating and operating public-cloud resources with Terraform or other infrastructure-as-code tools.
- Experience with Git, trunk-based development, feature flagging, Docker, Kubernetes orchestration, and end-to-end CI/CD pipelines.
- Experience participating in an unsupervised 24/7 on-call schedule and resolving issues without escalation.
- Experience with OpenTelemetry and monitoring tools such as Datadog, New Relic, Prometheus, Grafana, or Honeycomb.
- Understanding of networking and routing principles and security configuration for web/API services, including SSL and access control.
- Experience with backend database technologies, application containerization, reliability engineering, incident management, chaos engineering, load testing, or reliability testing.
- Experience with Jira or similar work-tracking systems and Confluence or similar documentation tools.
- Familiarity with security compliance frameworks including FedRAMP, NIST, and SOC 2.
- Ability to work with MacOS and predominantly command-line interfaces, understand stakeholder needs and business value, and communicate technical direction clearly.
