NOCD

Senior Site Reliability Engineer

NOCD
Apply
29 days ago

Base Salary

$160k - $200k/yr

Responsibilities

  • Build and maintain reliable production systems and services using production-quality code.
  • Develop APIs, microservices, automation, internal engineering tools, and platform capabilities.
  • Design, build, and operate AWS infrastructure using Terraform, Docker, and Kubernetes.
  • Improve CI/CD pipelines, deployment automation, release processes, and developer productivity.
  • Own reliability, availability, performance, scalability, monitoring, alerting, logging, and observability strategies.
  • Lead incident response, troubleshooting, root-cause analysis, postmortems, and preventative remediation.
  • Define and monitor SLIs, SLOs, operational metrics, and reliability practices.
  • Implement infrastructure and application security controls involving access management, secrets, encryption, logging, and compliance.
  • Own technical initiatives from design through production and collaborate across Software Engineering, Product, Security, and other teams.
  • Mentor engineers and help establish engineering and operational best practices.

Requirements

  • 7+ years of professional software engineering experience, including significant experience in SRE, platform engineering, DevOps, or cloud infrastructure.
  • Bachelor’s degree in Computer Science, Computer Engineering, Software Engineering, or a related technical field, or equivalent professional experience.
  • Strong software engineering fundamentals and experience writing production-quality code.
  • Strong hands-on experience with AWS, cloud architecture, Terraform or other infrastructure-as-code tools, Docker, and Kubernetes.
  • Proficiency in Python, TypeScript, or a similar programming language.
  • Experience designing and operating CI/CD pipelines and deployment automation.
  • Strong understanding of distributed systems, APIs, networking, databases, cloud architecture, reliability, scalability, availability, and performance.
  • Experience with monitoring, logging, observability, production troubleshooting, incident response, and root-cause analysis.
  • Ability to own technical initiatives and work effectively across engineering teams.
  • Preferred experience includes regulated environments, HIPAA, SOC 2, HITRUST, observability platforms, GitHub Actions, Jenkins, ArgoCD, GitOps, highly available or multi-region AWS architectures, microservices, event-driven architectures, DevSecOps, IAM, secrets management, encryption, SLIs, SLOs, SLAs, error budgets, developer platforms, disaster recovery, capacity planning, performance engineering, and high-growth startups.

Benefits

  • Comprehensive medical, dental, and vision coverage.
  • 401(k) match.
  • 11 observed company holidays per year.
  • PTO based on an accrual system.
  • Chicago office with an on-site gym.
  • 12 weeks of fully paid parental leave for the primary caregiver and 6 weeks for the secondary caregiver for qualifying full-time employees.
  • Hybrid work arrangement in Chicago with three days per week in the office.

Tech Stack

Argo CDAWSDatadogDockerGitHub ActionsGrafanaJenkinsKubernetesPrometheusPythonSplunkTerraformTypeScript

Categories

DevOpsSite Reliability
NOCD

About NOCD

501-1,000 employees

NOCD builds a telehealth platform for people with obsessive-compulsive disorder (OCD), combining live video sessions with licensed ERP therapists, self-guided tools, and peer support in a mobile app. It generates revenue by delivering virtual therapy and ongoing digital care through a national clinician network. Founded in 2018 and headquartered in Chicago, the company is privately held.

Contact me