
Infrastructure Engineer
Bayesian Health, Inc.4 months ago
Remote, United StatesSenior
Responsibilities
- Design cost-optimized, fault-tolerant AWS infrastructure that supports platform scale, new products, client deployments, and cloud cost management.
- Define branching and promotion strategies and build and maintain GitHub-based CI/CD pipelines for automated testing and deployment.
- Create infrastructure standards, guidelines, templates, and Terraform modules, and educate engineering and data science teams on their use.
- Monitor system performance and reliability, apply software upgrades, troubleshoot infrastructure issues, and optimize platform performance.
- Implement infrastructure security practices with the SecOps engineer while complying with HIPAA, HITRUST, FDA, and client requirements.
- Architect and build a secure internal AI Ops platform for hosting and managing AI/ML agents for infrastructure and DevOps optimization.
Requirements
- 5+ years of experience building and operating production cloud infrastructure on AWS in a DevOps, Infrastructure, Site Reliability Engineering, or similar role.
- Proficiency with Kubernetes, preferably Amazon EKS, including cluster bootstrapping and day-two operations.
- Strong operational knowledge of relational databases such as PostgreSQL and MySQL, including backups, failover, and performance tuning.
- Deep expertise in Terraform or equivalent infrastructure-as-code tooling and scalable module design.
- Familiarity with observability tools, particularly Datadog.
- Experience building infrastructure that handles sensitive PHI/PII data.
- Knowledge of CI/CD pipelines, preferably CircleCI.
- Strong communication and cross-functional collaboration skills, including translating engineering and data science requirements into technical solutions.
- Experience working through ambiguity and uncertainty in a startup environment.
- Preferred qualifications include AI-agent infrastructure optimization, disaster recovery or business continuity, multi-account and multi-cluster topologies, regulated-industry experience, chaos engineering or game-day facilitation, and GitOps framework implementation and maintenance.