LangChain

Infrastructure Engineer, Database

LangChain
Apply
5 hours ago

Base Salary

$180k - $230k/yr

Responsibilities

  • Own SmithDB deployment and operations across cloud environments, including cluster lifecycle management, blue/green and rolling upgrades, and automated failover.
  • Build and maintain infrastructure tooling with Terraform, Kubernetes, Helm, or equivalent to provision, configure, and scale SmithDB nodes.
  • Own the Kubernetes infrastructure supporting multi-tenant, high-throughput, low-latency distributed database services.
  • Build deployment pipelines, rollout strategies, and infrastructure-as-code for the storage layer.
  • Drive reliability engineering through incident response, postmortems, SLOs, and disaster recovery.
  • Manage capacity planning and cost efficiency, including growth modeling, resource rightsizing, and automated handling of customer traffic spikes.
  • Build safe, tested, and fast CI/CD promotion pipelines for database infrastructure changes.
  • Partner with SmithDB internals engineers on production-ready infrastructure changes and low-risk rollouts.

Requirements

  • 5+ years of experience in infrastructure, platform engineering, or SRE with hands-on experience.
  • Strong hands-on experience with Kubernetes and cloud infrastructure, including AWS, GCP, or Azure.
  • Scripting or systems programming ability in Go, Python, or a similar language.
  • Experience with infrastructure-as-code and CI/CD tooling such as Terraform, Helm, or ArgoCD.
  • Deep familiarity with at least one major cloud provider and the primitives used to run stateful workloads reliably.
  • Fluency with infrastructure-as-code, with Terraform or an equivalent such as Pulumi or CDK as a primary tool.
  • Experience operating high-traffic data systems, participating in on-call response, triaging incidents, and writing runbooks.
  • Experience with Kubernetes container orchestration and deploying stateful workloads in production.
  • Strong written and oral communication skills and the ability to explain infrastructure health to product and business stakeholders.
  • Experience owning production database systems such as Postgres, ClickHouse, Redis, or similar is preferred.
  • Comfort reading and reasoning about Rust is preferred.
  • Understanding of replication, backups, point-in-time recovery, connection pooling, and graceful degradation under load is preferred.

Benefits

  • Medical, dental, and vision coverage.
  • Flexible vacation.
  • 401(k) plan.
  • Meals on in-office days in the US.
  • Compensation includes base salary, variable compensation for relevant roles, equity, benefits, and perks.
  • Team members in the EU, UK, and APAC receive locally competitive benefits aligned with regional norms and regulations.

Categories

LangChain

About LangChain

201-500 employees

At LangChain, our mission is to make intelligent agents ubiquitous. We build the foundation for agent engineering in the real world, helping developers move from prototypes to production-ready AI agents that teams can rely on. What began as widely adopted open-source tools has grown into a platform for building, evaluating, deploying, and operating agents at scale. LangChain provides the agent engineering platform and open source frameworks developers need to ship reliable agents fast. LangSmith offers observability, evaluation, and deployment for rapid iteration. Our open source frameworks, LangGraph, LangChain, and Deep Agents, help developers build agents with speed and granular control. LangSmith is trusted by leading AI teams at Zip, Vanta, Klarna, Workday, Linkedin, Cloudflare, and more.