4 hours ago
Base Salary
$180k - $400k/yr
Responsibilities
- Build SLOs, alert routing, on-call operations, incident response, and postmortems for AI infrastructure.
- Own Terraform modules, state, testing, policy-as-code, drift management, infrastructure blueprints, and per-environment installations across AWS, Azure, and GCP.
- Automate customer environment provisioning, including firewall exceptions, DNS delegations, deployment identities, tag policies, and certificate chains.
- Maintain parity across AWS, Azure, and GCP using managed Kubernetes and appropriate cloud-native or portable services.
- Embed HIPAA and SOC 2 controls into infrastructure and generate compliance evidence through deployment pipelines.
- Own the Grafana, Loki, Mimir, Tempo, Alloy, Langfuse, and LiteLLM observability stack for operational analysis of autonomous agents.
- Create paved deployment paths for engineers shipping software into customer tenants.
- Participate in customer architecture reviews and coordinate with customer network teams.
- Investigate infrastructure patterns for autonomous AI systems, including agent observability and blast-radius controls.
Requirements
- 5+ years of experience operating production infrastructure in SRE, platform, or production engineering roles.
- Experience operating software in BYOC, single-tenant, on-premises, or air-gapped environments.
- Production experience with AWS, Azure, and GCP, including networking, IAM, and managed Kubernetes such as EKS, AKS, and GKE.
- Experience operating Kubernetes across multiple cloud providers.
- Deep Terraform experience with module design, state, testing, drift management, and blast-radius judgment.
- Experience with paging, incident command, postmortems, and SLOs.
- Experience operating under regulated requirements such as HIPAA, SOC 2, PCI, or FedRAMP.
- Ability to participate in customer-facing architecture reviews and infrastructure discussions.
- Programming or automation experience with Python, Go, or Bash.
- Nice-to-have experience with vendor control planes such as Ryvn, Nuon, or Replicated.
- Nice-to-have experience with GitOps, progressive delivery, multi-region infrastructure, and the Grafana stack at multi-tenant scale.
- Nice-to-have GPU and inference operations experience with Ray, SkyPilot, vLLM, or SGLang, as well as MLOps or research-to-production handoffs.
Benefits
- Competitive compensation, equity, 401(k) matching, comprehensive health and family benefits, and lunch every day.
- Office-first work arrangement emphasizing in-person collaboration.
- Culture centered on autonomy, trust, inclusion, and belonging.
Tech Stack
Categories
DevOpsSite Reliability
About Percepta
Percepta (a GC Transformation Company) combines applied AI engineering with frontier research to transform enterprises. Unlike traditional AI point solutions or consulting engagements, Percepta embeds AI engineers, researchers, and product managers directly within organizations. They are enabled by products from across GC’s portfolio and beyond, and leverage their own Mosaic platform to orchestrate enterprise transformation. Percepta exemplifies GC’s approach to deploying integrated, AI-native operating systems at scale.
