Staff Software Engineer, Platform Infrastructure
Pantheon Systems, Inc11 hours ago
Remote, IrelandStaff+
Responsibilities
- Establish SRE practices, including SLO/SLI frameworks, reliability standards, error budgets, toil reduction, and incident response, across the platform organization.
- Own observability patterns using Prometheus, OpenTelemetry, structured logging, and Grafana to provide actionable reliability signals.
- Build service templates, Terraform modules, Go APIs, CLIs, Cloud Run services, and GKE workloads for a self-service developer platform.
- Lead adoption of GCP-managed services including GCP Secret Manager, Cloud SQL, Cloud Run, and GKE, while standardizing GitHub Actions and Cloud Build.
- Apply security and compliance practices to networking, identity, secrets management, and platform architecture across SOC 2, ISO 27001, PCI DSS, and similar frameworks.
- Set technical direction, mentor engineers, influence adjacent teams, and drive reusable platform patterns.
- Own development, testing, operations, and support for platform systems in a full DevOps model.
- Participate in the on-call rotation after an initial ramp period of typically 3–6 months and reduce operational burden through automation, runbooks, and reliability improvements.
Requirements
- 8+ years of experience building and operating production systems, with significant infrastructure, SRE, or platform engineering experience.
- Demonstrated experience implementing SRE practices such as SLO frameworks, incident management, on-call culture, and measurable reliability improvements.
- Strong hands-on Go experience, including Go 1.21 or later, with Python as a secondary language.
- Deep practical experience with GCP services including Cloud Run, GKE, IAM, networking, and GCP-native managed services; comparable cloud experience may be considered with willingness to ramp on GCP.
- Hands-on experience designing and maintaining reusable Terraform modules.
- Solid operational experience with GKE or equivalent Kubernetes platforms.
- Production experience with Grafana, Prometheus, and OpenTelemetry across metrics, traces, and logs.
- Practical experience with IAP, IAM, VPC design, network controls, secrets management, and compliance frameworks such as SOC 2 or PCI DSS.
- Experience setting technical direction, translating ambiguous reliability goals into architecture, and influencing engineers without direct reporting relationships.
- Experience building reusable infrastructure patterns or developer platforms rather than one-off solutions.
- Clear communication skills for explaining reliability risks, architectural decisions, and incident postmortems to technical and non-technical stakeholders.
Benefits
- Industry-competitive compensation and an equity plan.
- 28 days of holiday and private medical and dental coverage.
- Life and critical illness insurance and a workplace pension scheme.
- Employee Resource Platform access, top-of-line equipment, and a monthly allowance for wellness and reading.
- Access to LinkedIn Learning and team-based and company-wide events and activities.
- Remote work within Ireland, with collaboration across North American and European time zones and flexibility for standups and incident response.
Tech Stack
DrupalGitHub ActionsGoGoogle Cloud PlatformGrafanaKubernetesNext.jsPrometheusPythonTerraformWordPress
Categories
DevOpsSite Reliability
About Pantheon Systems, Inc
Pantheon Systems provides a WebOps PaaS for building, hosting, and scaling Drupal, WordPress, and Next.js sites, with integrated CI/CD, governance, CDN, and security. It sells subscription plans to enterprises, agencies, and universities that manage multi-site portfolios and developer workflows in the cloud. Founded in 2010 and headquartered in San Francisco, the privately held company powers 300,000+ websites for organizations such as Google, Princeton, Clorox, Salesloft, and the United Nations.