7 hours ago
Base Salary
$180k - $400k/yr
Responsibilities
- Define infrastructure patterns for observable, controllable, and recoverable multi-agent systems.
- Own and evolve Terraform and Kubernetes infrastructure across AWS, GCP, and Azure.
- Build observability primitives that trace agent decisions and execution paths.
- Design and maintain CI/CD pipelines from code commit through production.
- Build operational foundations for monitoring, alerting, incident response, and AI-aware reliability practices.
- Collaborate across engineering teams to meet reliability and compliance requirements in regulated environments.
Requirements
- Require 5+ years of experience building and operating production infrastructure in DevOps or SRE roles.
- Require strong hands-on Terraform experience and deep experience with at least one major cloud provider: AWS, GCP, or Azure.
- Require solid production experience with Docker and Kubernetes, including managed clusters.
- Require experience designing and maintaining CI/CD pipelines using GitHub Actions, GitLab CI, or similar tools.
- Require scripting proficiency in Python, Bash, or a similar language.
- Prefer multi-region and multi-cloud experience across two or more providers.
- Prefer experience with single-tenant, on-premises, and multi-tenant SaaS deployments.
- Prefer familiarity with GitOps, progressive delivery, and the Grafana stack, including Prometheus, Grafana, and Loki.
- Prefer experience with HIPAA, SOC 2, and regulated infrastructure environments.
- Prefer experience supporting ML or research workflows moving into production, including model deployment or pipeline orchestration.
- Seek engineers with strong automation, communication, collaboration, agency, and curiosity about AI systems.
Benefits
- The role offers substantial ownership on a small team and the opportunity to define infrastructure patterns for autonomous AI systems.
- The company describes the work as mission-driven and fulfilling, with meaningful customer-facing problems.
- The role is not a 9-to-5 position and emphasizes flexibility around monitoring hours while expecting significant commitment to the work.
Tech Stack
AWSAzureBashDockerGitHub ActionsGitLab CI/CDGoogle Cloud PlatformGrafanaKubernetesPrometheusPythonTerraform
Categories
DevOpsSite Reliability
About Percepta
Percepta (a GC Transformation Company) combines applied AI engineering with frontier research to transform enterprises. Unlike traditional AI point solutions or consulting engagements, Percepta embeds AI engineers, researchers, and product managers directly within organizations. They are enabled by products from across GC’s portfolio and beyond, and leverage their own Mosaic platform to orchestrate enterprise transformation. Percepta exemplifies GC’s approach to deploying integrated, AI-native operating systems at scale.
