
Senior Site Reliability Engineer
Formation Bio2 days ago
Responsibilities
- Own the infrastructure and operational platform for shared engineering workloads, including compute, runtime environments, orchestration, deployment, observability, access controls, and reliability.
- Build and operate secure, observable, reliable infrastructure for product applications, containerized services, internal tools, data systems, ML pipelines, inference, and agentic software.
- Develop and maintain AWS infrastructure and additional cloud environments across development, staging, and production, including compute, networking, databases, load balancers, and secrets management.
- Create, review, maintain, and optimize infrastructure as code, CI/CD pipelines, and reusable platform patterns.
- Establish SLOs, monitoring, alerting, runbooks, incident response, advanced diagnostics, root cause analysis, and post-incident follow-through.
- Partner with Product Engineering, Data Engineering, and Data Science to implement architecture and infrastructure for product software, data systems, model training, and inference.
- Use AI tools to accelerate infrastructure development, investigate incidents, improve documentation, build automation, and make operational improvements while validating their output.
- Participate in support rotations and incident response, and mentor engineers on infrastructure and SRE fundamentals.
Requirements
- 5+ years of relevant experience in Site Reliability Engineering, infrastructure, systems, DevOps, or a similar discipline.
- Production experience operating cloud infrastructure and distributed systems with strong operational and reliability judgment.
- Experience with advanced diagnostics, incident response, root cause analysis, observability, and automation.
- Experience with AWS and Snowflake.
- Working experience with Docker, GitHub, Kubernetes, Python, Terraform or OpenTofu, and virtual networking.
- Experience managing shared commercial off-the-shelf and free/open-source software applications in addition to internal software and tools.
- Daily fluency with AI tools, including LLMs and agentic coding systems, with strong engineering judgment and validation standards.
- Experience supporting production ML or AI workloads, MLOps infrastructure, workflow orchestration, model serving, or related platforms is preferred.
- Experience with Azure, GCP, Vercel, or Terragrunt is preferred.
- Experience operating infrastructure in a regulated or validated environment is preferred.
- Strong collaboration and communication skills across technical and non-technical partners.
Benefits
- The posting states that equity, comprehensive benefits, and generous perks are offered in addition to base salary.
- Hybrid work requires three days per week in the office.
- Hiring is prioritized in the New York City and Boston metro areas; applicants in the Research Triangle and San Francisco Bay Area may also be considered, and applicants must reside in or be willing to relocate to these locations.
Tech Stack
Categories
DevOpsSite Reliability
About Formation Bio
Formation Bio builds AI-driven technology platforms and runs a pharmaceutical development business that designs and executes clinical trials for in-licensed or acquired drugs. It partners with pharma, biotechs, and research groups to take clinical-stage assets through proof of concept and beyond, aiming to reduce trial time and improve data quality. The privately held company is headquartered in New York and originated as TrialSpark before rebranding to Formation Bio.