7 months ago
San Francisco, CA, USAMid Level
Base Salary
$175k - $250k/yr
Responsibilities
- Architect, operate, and improve the infrastructure powering the 25ms cold-start compute engine.
- Build and evolve observability across metrics, traces, and logs.
- Define, monitor, and drive SLOs and SLIs across critical system surfaces.
- Lead incident response, root cause analysis, post-mortems, and systemic remediation.
- Design self-healing and automated operational systems to reduce toil and scale operations.
- Tune performance across compute, networking, storage, and sandboxed execution layers.
- Build automation and tooling for operations, debugging, capacity planning, and failure prediction.
- Conduct load testing, chaos engineering, and performance benchmarking.
- Own infrastructure-layer security practices, including sandboxed compute and network isolation.
- Partner with platform engineers to incorporate reliability into new features.
Requirements
- At least 3 years of experience in SRE, DevOps, or infrastructure engineering roles.
- Strong proficiency in at least one programming language, such as Go, Rust, or Python.
- Hands-on experience with a major cloud provider, including AWS or GCP.
- Knowledge of Linux systems, networking fundamentals, and distributed systems.
- Experience with bare-metal servers and datacenter operations, including PXE/iPXE provisioning, IPMI/BMC, RAID/NVMe, SR-IOV, and high-throughput networking.
- Experience with Kubernetes or similar orchestrators.
- Familiarity with observability stacks such as Prometheus, Grafana, ELK, or Datadog.
- Experience building and maintaining CI/CD pipelines using GitHub Actions, GitLab CI, or Jenkins.
- Strong debugging, problem-solving, and incident-management skills.
- Preferred experience includes Terraform or Pulumi, service mesh or API gateway technologies, chaos engineering or resiliency testing, cloud security, and high-growth or high-availability environments.
- Bonus experience includes serverless compute, sandboxed execution environments, ultra-low-latency runtime engineering, distributed key-value stores and databases, chaos engineering, systems-level programming, and generative AI infrastructure.
Tech Stack
AWSDatadogGitHub ActionsGitLab CI/CDGoGoogle Cloud PlatformGrafanaJenkinsKubernetesLinuxPrometheusPythonRustTerraform
Categories
DevOpsSite Reliability
About Blaxel
Blaxel keeps infinite, secure sandboxes on automatic standby to run AI code, while co-hosting your agents for near instant latency.
