Airbyte

Senior Site Reliability Engineer - Hiring Sprint

Airbyte
Apply
2 months ago

Base Salary

$196k - $255k/yr

Responsibilities

  • Own the infrastructure supporting the Data Replication platform across Kubernetes clusters, CI/CD pipelines, secrets management, networking, and cloud resources in AWS and GCP
  • Partner with product engineers to integrate product features reliably with infrastructure
  • Maintain and improve observability, alerting, and anomaly detection with potential LLM automation
  • Build and maintain AI-augmented release and internal tooling, including canary deployments, progressive rollouts, automated release qualification, and rollback automation
  • Set infrastructure standards, build self-service tooling, write runbooks, and coach engineers to own more of their stack
  • Drive down incidents and establish reliability standards for the team

Requirements

  • 7+ years of experience in infrastructure, platform engineering, SRE, or DevOps
  • Hands-on production ownership of Kubernetes, Helm, and Terraform
  • Deep experience with observability stacks including Prometheus, Grafana, and Datadog
  • Experience owning CI/CD pipelines and developer tooling
  • Ability and willingness to read backend code to understand system failures and instrument systems correctly
  • Fluency with AI tools, LLMs, and agentic frameworks for automation, debugging, and reducing toil
  • Comfort with ambiguity, rapid execution, and end-to-end ownership
  • Experience with data pipelines, replication systems, or ETL/ELT platforms is preferred
  • Experience with control plane/data plane architectures or internal developer platforms is preferred
  • Experience with Airbyte, CDKs, or connector-based architectures is preferred

Benefits

  • Flexible PTO with a culture encouraging at least 25 days off annually
  • 16 weeks of fully paid parental leave for all parents
  • Comprehensive medical, dental, and vision coverage for employees and dependents
  • 401(k) retirement plan
  • Professional development budget, conference sponsorship, and book reimbursement
  • Commuter benefits and monthly internet reimbursement
  • Breakfast and lunch in the San Francisco office
  • Onsite four days per week in San Francisco, California

Tech Stack

Categories

DevOpsSite Reliability
Airbyte

About Airbyte

51-200 employees

Founded in 2020, Airbyte is the context layer for production-grade AI agents. AI agents fail in production for one reason: they can't see the business. They make scattered API calls at runtime, burn tokens reconciling fragmented data, and break under real workloads. Airbyte solves this by giving agents unified, permission-aware access to the operational data scattered across the tools companies actually run on (CRM, billing, support, product, internal systems) through a hybrid architecture that combines large-scale replication (for cross-system search and discovery) with real-time fetching (for fresh operational state). It's the combination production agents actually need, built on the connector footprint we've hardened over six years. We've raised $181M from Benchmark, Accel, Altimeter, Coatue, Y Combinator, and others. Today, 25,000+ companies sync data with Airbyte's 600+ connectors. Open source remains core to how we build, because the data foundation under your AI agents is too important to be a black box.