2 months ago
Responsibilities
- Design and evolve GCP cloud architecture, including networking, interconnects, IAM, and high-availability topology, using Terraform and GitOps.
- Build and own CI/CD pipelines for infrastructure-as-code with policy guardrails, drift detection, and progressive rollout.
- Develop self-service platform capabilities and golden paths for engineering teams.
- Improve observability across Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager.
- Operate GKE clusters, Helm-packaged workloads, RabbitMQ and IBM MQ message brokers, and data stores.
- Participate in Follow-The-Sun on-call, alert triage, incident response, structured debugging, escalation, and blameless post-mortems.
- Embed SRE practices such as SLIs, SLOs, error budgets, and capacity planning into infrastructure operations.
Requirements
- At least 5 years of experience in DevOps, platform/infrastructure, or SRE roles operating large-scale, highly available, high-performance production systems.
- Deep hands-on experience designing cloud architecture on Google Cloud Platform, including landing zones, networking, IAM, and high-availability topology.
- Strong Infrastructure-as-Code experience with Terraform across multiple environments, with GitOps and least-privilege practices.
- Experience building CI/CD pipelines for infrastructure-as-code, including automated plan/apply, code review, policy-as-code, drift detection, and safe rollout.
- Significant production experience with Kubernetes, ideally GKE, and Helm.
- Strong L3/L4-L7 networking fundamentals, including VPCs, routing, load balancing, DNS, TLS, and interconnects.
- Hands-on experience with Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager for metrics, logs, traces, and alerting.
- Operator-level familiarity with PostgreSQL and message brokers such as RabbitMQ and RedPanda.
- Understanding of SRE practices including SLOs, error budgets, and capacity planning.
- Strong incident-management skills covering structured debugging, escalation, documentation, and post-mortems.
- Willingness to participate in a Follow-The-Sun on-call rotation from APAC hours and work effectively in a distributed, async-first team.
- Bonus experience with OPA/Conftest, Checkov, tflint, Atlantis, Terraform state and module registries, Backstage, Tilt, Alloy collector, Rootly, Go, Linux, Docker, containerd, SOC 2, secrets management, audit logging, or regulated fintech systems.
Benefits
- Competitive salary and stock options.
- Health benefits.
- One-time USD $500 new-hire home-office setup.
- USD $150 monthly stipend via a Brex Card.
- Globally distributed, async-first work environment with a Follow-The-Sun on-call rotation from APAC hours.
Tech Stack
Categories
About Alpaca
Alpaca is a US-headquartered, self-clearing broker-dealer and a global leader in brokerage infrastructure APIs providing access to stocks, ETFs, options, fixed income, and crypto. Alpaca delivers embeddable finance solutions for tokenization, fully paid securities lending, high-yield cash, 24/5 trading, Shariah-compliant investing, and more. Today, Alpaca powers over 9 million brokerage accounts across hundreds of fintechs and institutions in 40+ countries, with over $320M in funding from top investors including Drive Capital, Portage Ventures, Spark Capital, Tribe Capital, Social Leverage, Derayah Financial, Horizons Ventures, Unbound, SBI Group, Elefund, and Y Combinator.
