1 day ago
Base Salary
$105k - $163k/yr
Responsibilities
- Monitor critical systems and maintain real-time dashboards to identify performance and availability anomalies.
- Participate in on-call incident response, troubleshoot issues, mitigate outages, and restore services.
- Investigate incidents, identify root causes, document findings, and collaborate on recurrence prevention.
- Improve monitoring, logging, alerting, and observability practices.
- Write and maintain scripts for operational automation, incident remediation, and reporting.
- Support the definition and tracking of Service Level Objectives and Service Level Indicators.
- Assist with CI/CD pipeline and workflow improvements to minimize downtime.
- Create and maintain documentation for monitoring configurations, incident procedures, and root cause analyses.
- Collaborate with software engineering and infrastructure teams to improve fault tolerance, scalability, and operational readiness.
Requirements
- Bachelor’s degree in computer science, engineering, or a related field, or equivalent practical experience.
- Foundational understanding of site reliability engineering principles, monitoring, alerting, and incident management.
- Exposure to observability tools including Prometheus, Datadog, New Relic, or Grafana and logging platforms such as Splunk or Elasticsearch.
- Basic proficiency in one or more programming or scripting languages, including Python, Go, Bash, or Java.
- Familiarity with AWS, Google Cloud Platform, or Azure and their services.
- Understanding of Docker and Kubernetes.
- Strong analytical, communication, documentation, collaboration, and problem-solving skills.
- Proactive mindset, attention to detail, and willingness to learn in a fast-paced environment.
Benefits
- Medical, vision, dental, retirement, paid time away, life insurance, and disability benefits.
- Merchandise discount and Employee Assistance Program resources.
- 401(k), insurance options, PTO accruals, holidays, and additional benefits subject to eligibility requirements.
- May be eligible for performance-based incentives or bonuses.
- Onsite in Seattle, Washington, five days per week, with a Sunday–Thursday schedule from 6:30 a.m. to 2:30 p.m.
Tech Stack
AWSAzureBashDatadogDockerElasticsearchGoGoogle Cloud PlatformGrafanaJavaKubernetesPrometheusPythonSplunk
Categories
Site Reliability
About Nordstrom
Nordstrom is a U.S. fashion retailer selling apparel, shoes, accessories, beauty, and home goods through department stores, off-price Nordstrom Rack, and e-commerce. Founded in 1901 and headquartered in Seattle, it is a public company (NYSE: JWN) serving consumers nationwide, with technology, merchandising, and supply-chain teams supporting its digital and in-store operations.
