7 days ago
Responsibilities
- Monitor production availability and take a holistic view of system health.
- Build software and systems to manage platform infrastructure and applications.
- Improve reliability, quality, performance, and time-to-market for software solutions.
- Provide primary operational support and engineering for multiple large distributed applications.
- Gather and analyze operating-system and application metrics for performance tuning and fault finding.
- Partner with development teams on testing and release procedures.
- Participate in system design consulting, platform management, and capacity planning.
- Create sustainable systems and services through automation and platform improvements.
- Balance feature delivery speed and reliability using defined service-level objectives.
- Drive incident response, resolution, and cross-functional communication during outages and critical incidents.
- Conduct blameless postmortems and support continuous reliability improvements.
Requirements
- 3–6 years of working experience in a similar role focused on systems engineering, automation, and reliability.
- Proficiency in at least one programming language such as Python, Go, Java, or C# and experience with Bash or PowerShell scripting.
- Deep understanding of cloud computing platforms, including AWS and services such as EC2, ECS, Lambda, and DynamoDB.
- Experience with infrastructure-as-code tools such as CloudFormation and Terraform.
- Deep understanding of CI/CD concepts and experience with Jenkins, GitLab CI/CD, or CircleCI.
- Strong knowledge of Docker, Kubernetes, and microservices architecture.
- Experience with monitoring and observability tools such as Prometheus, Grafana, ELK stack, and CloudWatch.
- Ability to troubleshoot complex issues in distributed systems.
- Experience with incident management, blameless postmortems, and driving incident response efforts.
- Preferred experience with large Kubernetes clusters and Grafana Observability Suite tools including Loki, Mimir, and Tempo.
- Preferred experience with Splunk, Datadog, PagerDuty, Rundeck, Ansible, Puppet, or Chef.
- Relevant certifications such as AWS Certified DevOps Engineer or Google Cloud Professional DevOps Engineer are advantageous.
- Strong communication, collaboration, prioritization, self-learning, and composure under pressure.
Benefits
- NICE-FLEX hybrid work model with two days in the office and three remote days each week.
- Individual contributor role reporting to the Director of Network Operations.
- Equal opportunity employment.
Tech Stack
Categories
Site Reliability
About NICE
NiCE is transforming the world with AI that puts people first. Our purpose-built AI-powered platforms automate engagements into proactive, safe, intelligent actions, empowering individuals and organizations to innovate and act, from interaction to resolution. Trusted by organizations throughout 150+ countries worldwide, NiCE’s platforms are widely adopted across industries connecting people, systems, and workflows to work smarter at scale, elevating performance across the organization, delivering proven measurable outcomes.
