DroneUp

SRE - Platform Engineer

DroneUp
Apply
8 months ago
Remote, United StatesStaff+

Responsibilities

  • Architect the internal developer platform and cloud engineering capabilities across internal and client-facing infrastructure.
  • Design and oversee GCP environments, GKE cluster operations, workload management, and public-cloud resources.
  • Define and maintain SLOs, SLIs, and error budgets while improving system reliability and availability.
  • Lead incident response, on-call rotations, root-cause analysis, and post-mortem reviews.
  • Implement and optimize monitoring, alerting, and observability systems using tools such as OpenTelemetry, Prometheus, Grafana, and Honeycomb.
  • Drive automation, self-service, security by default, least privilege, resilient systems, and infrastructure engineering practices.
  • Mentor platform engineers and enable platform engineering capabilities within software engineering teams.
  • Develop and maintain tooling, infrastructure, tests, documentation, runbooks, ADRs, and CI/CD pipelines.
  • Collaborate on capacity planning, performance optimization, peer reviews, pair programming, and continuous improvement.
  • Promote testing, chaos engineering, load testing, reliability testing, and strong engineering practices across the platform.

Requirements

  • Bachelor's degree in Computer Science, Computer Engineering, or a related field, or 8+ years of experience as a software engineer.
  • Proficiency with Kubernetes and extensive experience with Unix/Linux; CKA or CKAD certification is optional.
  • Proficiency in multiple languages or infrastructure languages, ideally including Golang, Node.js, Python, and HCL.
  • Familiarity with at least two of GCP, AWS, and Azure in multi-cloud environments.
  • Advanced experience creating and operating public-cloud resources with Terraform or other infrastructure-as-code tools.
  • Experience with Git, trunk-based development, feature flagging, Docker, Kubernetes orchestration, and end-to-end CI/CD pipelines.
  • Experience participating in an unsupervised 24/7 on-call schedule and resolving issues without escalation.
  • Experience with OpenTelemetry and monitoring tools such as Datadog, New Relic, Prometheus, Grafana, or Honeycomb.
  • Understanding of networking and routing principles and security configuration for web/API services, including SSL and access control.
  • Experience with backend database technologies, application containerization, reliability engineering, incident management, chaos engineering, load testing, or reliability testing.
  • Experience with Jira or similar work-tracking systems and Confluence or similar documentation tools.
  • Familiarity with security compliance frameworks including FedRAMP, NIST, and SOC 2.
  • Ability to work with MacOS and predominantly command-line interfaces, understand stakeholder needs and business value, and communicate technical direction clearly.

Tech Stack

Categories

DevOpsSite Reliability
Contact me