Base Salary
$196k - $255k/yr
Responsibilities
- Own the infrastructure supporting the Data Replication platform across Kubernetes clusters, CI/CD pipelines, secrets management, networking, and cloud resources in AWS and GCP
- Partner with product engineers to integrate product features reliably with infrastructure
- Maintain and improve observability, alerting, and anomaly detection with potential LLM automation
- Build and maintain AI-augmented release and internal tooling, including canary deployments, progressive rollouts, automated release qualification, and rollback automation
- Set infrastructure standards, build self-service tooling, write runbooks, and coach engineers to own more of their stack
- Drive down incidents and establish reliability standards for the team
Requirements
- 7+ years of experience in infrastructure, platform engineering, SRE, or DevOps
- Hands-on production ownership of Kubernetes, Helm, and Terraform
- Deep experience with observability stacks including Prometheus, Grafana, and Datadog
- Experience owning CI/CD pipelines and developer tooling
- Ability and willingness to read backend code to understand system failures and instrument systems correctly
- Fluency with AI tools, LLMs, and agentic frameworks for automation, debugging, and reducing toil
- Comfort with ambiguity, rapid execution, and end-to-end ownership
- Experience with data pipelines, replication systems, or ETL/ELT platforms is preferred
- Experience with control plane/data plane architectures or internal developer platforms is preferred
- Experience with Airbyte, CDKs, or connector-based architectures is preferred
Benefits
- Flexible PTO with a culture encouraging at least 25 days off annually
- 16 weeks of fully paid parental leave for all parents
- Comprehensive medical, dental, and vision coverage for employees and dependents
- 401(k) retirement plan
- Professional development budget, conference sponsorship, and book reimbursement
- Commuter benefits and monthly internet reimbursement
- Breakfast and lunch in the San Francisco office
- Onsite four days per week in San Francisco, California
Tech Stack
Categories
About Airbyte
Founded in 2020, Airbyte is the context layer for production-grade AI agents. AI agents fail in production for one reason: they can't see the business. They make scattered API calls at runtime, burn tokens reconciling fragmented data, and break under real workloads. Airbyte solves this by giving agents unified, permission-aware access to the operational data scattered across the tools companies actually run on (CRM, billing, support, product, internal systems) through a hybrid architecture that combines large-scale replication (for cross-system search and discovery) with real-time fetching (for fresh operational state). It's the combination production agents actually need, built on the connector footprint we've hardened over six years. We've raised $181M from Benchmark, Accel, Altimeter, Coatue, Y Combinator, and others. Today, 25,000+ companies sync data with Airbyte's 600+ connectors. Open source remains core to how we build, because the data foundation under your AI agents is too important to be a black box.
