3 months ago
Bucharest, RomaniaStaff+
Responsibilities
- Deploy, operate, and maintain blockchain RPC nodes across multiple chains and geographic regions.
- Manage Kubernetes clusters and the platform supporting blockchain nodes.
- Perform rolling upgrades and hard fork migrations for EVM and non-EVM blockchain clients.
- Participate in on-call rotations, triage incidents with PagerDuty, and coordinate resolution of outages, latency spikes, and SLO breaches.
- Develop and maintain AI agents and automation for health checks, auto-healing, and hard fork notifications.
- Deploy and manage services using ArgoCD, GitOps workflows, and Helm charts.
- Provision, benchmark, and maintain bare-metal and cloud infrastructure, including hardware replacement.
- Respond to security advisories and coordinate upgrades with minimal downtime.
- Contribute to postmortems, async reviews, action-item tracking, and resolution follow-up.
- Collaborate with product, customer success, and engineering teams on chain deprecations, capacity planning, and SLO reporting.
Requirements
- Experience designing and operating large-scale, multi-region, multi-cloud production systems.
- Experience with Kubernetes, including k3s or similar platforms, StatefulSets, storage management, Secrets, and service meshes such as Istio.
- Experience managing secrets and access control in multi-cluster environments.
- Familiarity with automation frameworks for node health checks, upgrades, and remediation workflows.
- Experience with Infrastructure-as-Code tools such as Terraform, Ansible, Pulumi, CloudFormation, Chef, or Puppet.
- Experience with GitOps tooling, including ArgoCD and Helm.
- Experience with cloud infrastructure, bare-metal management, storage provisioning, and snapshot management.
- Proficiency with Grafana, Prometheus, and Alertmanager, including building or tuning dashboards and alert rules.
- Comfort working in an on-call environment and triaging production incidents using PagerDuty and structured runbooks.
- Ability to write technical documentation and postmortems and contribute to async-first communication.
- Experience with networking and configuring or managing VPC networks.
- Basic understanding of security best practices.
- Preferred experience with service mesh deployments such as Istio, web applications, microservice architecture, startups, and blockchain or Web3 technologies.
Benefits
- Attractive salary package.
- Flexible time away.
- Private medical insurance.
- Internal off-site hackathons.
- Access to a company-rented hacker house during summer.
- Opportunity to travel across offices.