1 day ago
Base Salary
$95k - $159k/yr
Responsibilities
- Create monitoring queries and establish service-level baselines.
- Support senior engineers during incidents and contribute to post-mortems and root cause analyses.
- Participate in disaster recovery tests and test service availability, reliability, and recoverability in non-production environments.
- Implement automation and execute code in production environments.
- Document SRE knowledge and support the deployment, monitoring, reliability, and security of services integrating AI tools.
- Support infrastructure topology drawings, deployment workflows, service handover, and capability building.
- Support multiple engineering teams by troubleshooting infrastructure and application layers and enabling secure self-service practices.
Requirements
- Expertise in advanced Terraform, including modules, providers, state management, lifecycle controls, drift detection, safe refactoring, remote state, locking, and cross-stack dependencies.
- Hands-on experience managing production, multi-account, multi-region AWS environments using ECS, RDS, ALB, VPC, IAM, Route 53, ECR, S3, Lambda, DynamoDB, SQS, Secrets Manager, KMS, and CloudWatch.
- Experience building and troubleshooting reusable GitHub Actions workflows with OIDC authentication, approval gates, runners, Terraform deployments, application deployments, and migration pipelines.
- Knowledge of Docker, ECR, ECS task definitions and services, IAM roles, health checks, autoscaling, ALB integration, and deployment rollbacks.
- Proficiency in AWS networking and security, including VPCs, ALBs, Route 53, ACM/TLS, IAM, OIDC, Secrets Manager, KMS, and cloud security practices.
- Skill in incident response and observability using logs, metrics, alarms, deployment history, root cause analysis, rollback decisions, and operational runbooks.
- Strong Linux and Git fundamentals with Bash and Python scripting for AWS CLI automation, CI/CD, and operational tooling.
- Hands-on experience integrating and operating AI services and APIs in production, including monitoring, reliability, and security practices.
- Ability to support engineering teams, troubleshoot across infrastructure and application layers, document solutions, and enable secure self-service practices.
Benefits
- U.S. national base pay range of $95,300-$158,800, with geographic differentials potentially applying; compensation figures excluded from structured benefits.
- Eligible for an annual incentive bonus.
- Country-specific benefits are available.
- Accommodation support is available for candidates with disabilities or other needs.
- The employer provides an equal opportunity hiring process.
Categories
DevOpsSite Reliability
About Elsevier
Elsevier builds scientific and medical information products and analytics platforms for researchers, clinicians, and educators. Its portfolio includes journals and databases such as The Lancet, Cell, ScienceDirect, Scopus, and ClinicalKey, sold primarily via institutional subscriptions, licensing, and open-access publishing fees. Founded in 1880 and headquartered in Amsterdam, Elsevier operates as a business of RELX, serving universities, hospitals, corporations, and government research agencies worldwide.
