Blaxel

Site Reliability Engineer

Blaxel
Apply
7 months ago

Base Salary

$175k - $250k/yr

Responsibilities

  • Architect, operate, and improve the infrastructure powering the 25ms cold-start compute engine.
  • Build and evolve observability across metrics, traces, and logs.
  • Define, monitor, and drive SLOs and SLIs across critical system surfaces.
  • Lead incident response, root cause analysis, post-mortems, and systemic remediation.
  • Design self-healing and automated operational systems to reduce toil and scale operations.
  • Tune performance across compute, networking, storage, and sandboxed execution layers.
  • Build automation and tooling for operations, debugging, capacity planning, and failure prediction.
  • Conduct load testing, chaos engineering, and performance benchmarking.
  • Own infrastructure-layer security practices, including sandboxed compute and network isolation.
  • Partner with platform engineers to incorporate reliability into new features.

Requirements

  • At least 3 years of experience in SRE, DevOps, or infrastructure engineering roles.
  • Strong proficiency in at least one programming language, such as Go, Rust, or Python.
  • Hands-on experience with a major cloud provider, including AWS or GCP.
  • Knowledge of Linux systems, networking fundamentals, and distributed systems.
  • Experience with bare-metal servers and datacenter operations, including PXE/iPXE provisioning, IPMI/BMC, RAID/NVMe, SR-IOV, and high-throughput networking.
  • Experience with Kubernetes or similar orchestrators.
  • Familiarity with observability stacks such as Prometheus, Grafana, ELK, or Datadog.
  • Experience building and maintaining CI/CD pipelines using GitHub Actions, GitLab CI, or Jenkins.
  • Strong debugging, problem-solving, and incident-management skills.
  • Preferred experience includes Terraform or Pulumi, service mesh or API gateway technologies, chaos engineering or resiliency testing, cloud security, and high-growth or high-availability environments.
  • Bonus experience includes serverless compute, sandboxed execution environments, ultra-low-latency runtime engineering, distributed key-value stores and databases, chaos engineering, systems-level programming, and generative AI infrastructure.

Tech Stack

AWSDatadogGitHub ActionsGitLab CI/CDGoGoogle Cloud PlatformGrafanaJenkinsKubernetesLinuxPrometheusPythonRustTerraform

Categories

DevOpsSite Reliability
Blaxel

About Blaxel

1-10 employees

Blaxel keeps infinite, secure sandboxes on automatic standby to run AI code, while co-hosting your agents for near instant latency.

Contact me