10 days ago
Remote, United StatesSenior
Base Salary
$160k - $190k/yr
Responsibilities
- Own core platform services and major infrastructure migrations end to end, including stateful and business-critical systems.
- Operate containerized workloads and delivery systems using orchestration, GitOps-style deployment, progressive rollouts, rollback strategies, and service mesh technologies.
- Architect and operate AWS foundations including multi-account governance, IAM, networking, DNS, certificates, private connectivity, and egress paths.
- Manage workload identity, least privilege, SSO, OIDC, machine-to-machine credentials, Vault, encryption, and key rotation.
- Operate databases, message brokers, caches, time-series stores, backups, replication, upgrades, and restores.
- Build observability through distributed tracing, metrics, SLOs, alerting, dashboards, and testing frameworks for dispatch and market-message workflows.
- Develop infrastructure-as-code and internal developer tooling using Terraform, GitHub, Buildkite, Docker, Nomad, and internal tools.
- Build secure infrastructure and guardrails for AI-assisted development and access to internal systems.
- Participate in on-call ownership, mentor engineers, document systems, and coordinate shared infrastructure changes across teams.
Requirements
- Strong production software development experience in Go and/or Python, including maintaining services and tooling and writing tests.
- Approximately 6+ years of professional software engineering experience, including several years in DevOps or SRE operating production systems.
- Demonstrated ownership of infrastructure projects and migrations involving stateful or business-critical systems with minimal disruption.
- Deep production AWS experience with multi-account organizations, IAM, cross-account access, VPC, network design, DNS, secrets management, and encryption key management.
- Deep hands-on Kubernetes production experience, including cluster upgrades, networking, ingress, RBAC, resource management, autoscaling, and debugging workloads under load.
- Strong infrastructure-as-code experience with Terraform or a similar tool.
- Strong monitoring and observability experience, including metrics, alerts, dashboards, distributed tracing, and open-source tools such as Prometheus, Grafana, Elasticsearch, or OpenSearch.
- Hands-on experience operating stateful systems such as relational databases and message brokers, including upgrades, replication changes, or restores.
- Ability to operate and improve systems built by others, communicate clearly, document work, mentor teammates, and coordinate across teams.
- Excitement about AI-assisted development and experience or interest in tools such as Claude Code, MCP, and agents.
- Preferred experience includes Nomad, Consul, Vault, EKS, ArgoCD, Flux, Helm, Jenkins, OpenTelemetry, MSK, AWS Organizations, Control Tower, Auth0, Okta, Cognito, Keycloak, internal developer tooling, AWS Bedrock, or running model infrastructure.
- Preferred ability to read and modify Java or C++ services and interest in the energy industry and clean-energy transition.
- Applicants must include a link to their GitHub account; applications without one will not be considered.
Benefits
- Fully remote role with employees distributed across the US, Canada, and abroad.
- Must overlap with the Eastern time zone for at least four hours of the workday.
- Employees must generally work from their home country, and international work travel requires approval under the Global Remote Travel Policy.
- The company does not sponsor visas or transfers for new hires; employees must be authorized to work from their home location.
