1 day ago
Base Salary
$86k - $127k/yr
Responsibilities
- Design, build, and operate resilient and scalable systems using DevOps, SRE, and software engineering practices.
- Deliver high-availability services through automation, infrastructure as code, and reliability engineering.
- Improve monitoring, logging, alerting, and observability for distributed systems.
- Support CI/CD automation, deployment workflows, and production tooling to reduce operational toil.
- Lead incident response, root cause analysis, and recovery improvements to minimize downtime.
- Partner with engineering teams to embed reliability into the software development lifecycle.
- Automate provisioning, configuration, and self-healing across cloud and on-premises environments.
- Validate resiliency and performance through testing, chaos engineering, and capacity planning.
Requirements
- At least 9 years of experience developing and designing products around Site Reliability Engineering principles for containerized workloads and on-premises services using Kubernetes.
- Experience managing and interpreting large datasets, querying data, and creating dashboards and reports with Power BI and Grafana.
- Experience managing cloud and on-premises infrastructure with infrastructure as code tools including Terraform and CloudFormation.
- Hands-on experience building, operating, monitoring, logging, and alerting distributed systems at scale using Datadog and Splunk.
- Experience supporting service delivery and operations with Jenkins, Azure DevOps, Team Foundation Version Control, and CI/CD automation.
- Experience developing software and automation solutions using Python.
- Knowledge of scalability and resiliency practices for applications deployed on AWS and Azure, including Lambda and API Gateway.
- Strong scripting, automation, and integration experience across Linux and Windows environments.
Benefits
- Full-time position located in Boston, Massachusetts.
- Base compensation range of $86129-$127189, with possible additional variable compensation, bonuses, or commissions.
- Paid time off, including vacation, company holidays, personal days, and sick leave, subject to employee grade and policy.
- Medical, dental, and vision coverage, or provincial healthcare coordination in Canada.
- Retirement savings plans such as 401(k) in the U.S. and RRSP in Canada.
- Life and disability insurance and employee assistance programs.
Categories
DevOpsSite Reliability
About Capgemini
Capgemini is a global IT services and consulting firm that delivers strategy, cloud, AI, software engineering, and managed services to large enterprises and public-sector clients. Founded in 1967 and headquartered in Paris, it is publicly traded on Euronext Paris and operates in 50+ countries. The group expanded its engineering capabilities by acquiring Altran in 2020, now operating as Capgemini Engineering.
