29 days ago
Base Salary
$160k - $200k/yr
Responsibilities
- Build and maintain reliable production systems and services using production-quality code.
- Develop APIs, microservices, automation, internal engineering tools, and platform capabilities.
- Design, build, and operate AWS infrastructure using Terraform, Docker, and Kubernetes.
- Improve CI/CD pipelines, deployment automation, release processes, and developer productivity.
- Own reliability, availability, performance, scalability, monitoring, alerting, logging, and observability strategies.
- Lead incident response, troubleshooting, root-cause analysis, postmortems, and preventative remediation.
- Define and monitor SLIs, SLOs, operational metrics, and reliability practices.
- Implement infrastructure and application security controls involving access management, secrets, encryption, logging, and compliance.
- Own technical initiatives from design through production and collaborate across Software Engineering, Product, Security, and other teams.
- Mentor engineers and help establish engineering and operational best practices.
Requirements
- 7+ years of professional software engineering experience, including significant experience in SRE, platform engineering, DevOps, or cloud infrastructure.
- Bachelor’s degree in Computer Science, Computer Engineering, Software Engineering, or a related technical field, or equivalent professional experience.
- Strong software engineering fundamentals and experience writing production-quality code.
- Strong hands-on experience with AWS, cloud architecture, Terraform or other infrastructure-as-code tools, Docker, and Kubernetes.
- Proficiency in Python, TypeScript, or a similar programming language.
- Experience designing and operating CI/CD pipelines and deployment automation.
- Strong understanding of distributed systems, APIs, networking, databases, cloud architecture, reliability, scalability, availability, and performance.
- Experience with monitoring, logging, observability, production troubleshooting, incident response, and root-cause analysis.
- Ability to own technical initiatives and work effectively across engineering teams.
- Preferred experience includes regulated environments, HIPAA, SOC 2, HITRUST, observability platforms, GitHub Actions, Jenkins, ArgoCD, GitOps, highly available or multi-region AWS architectures, microservices, event-driven architectures, DevSecOps, IAM, secrets management, encryption, SLIs, SLOs, SLAs, error budgets, developer platforms, disaster recovery, capacity planning, performance engineering, and high-growth startups.
Benefits
- Comprehensive medical, dental, and vision coverage.
- 401(k) match.
- 11 observed company holidays per year.
- PTO based on an accrual system.
- Chicago office with an on-site gym.
- 12 weeks of fully paid parental leave for the primary caregiver and 6 weeks for the secondary caregiver for qualifying full-time employees.
- Hybrid work arrangement in Chicago with three days per week in the office.
Tech Stack
Argo CDAWSDatadogDockerGitHub ActionsGrafanaJenkinsKubernetesPrometheusPythonSplunkTerraformTypeScript
Categories
DevOpsSite Reliability
About NOCD
NOCD builds a telehealth platform for people with obsessive-compulsive disorder (OCD), combining live video sessions with licensed ERP therapists, self-guided tools, and peer support in a mobile app. It generates revenue by delivering virtual therapy and ongoing digital care through a national clinician network. Founded in 2018 and headquartered in Chicago, the company is privately held.
