4 hours ago
Responsibilities
- Monitor production availability and maintain a holistic view of system health.
- Build software and systems to manage platform infrastructure and applications.
- Improve the reliability, quality, performance, and time-to-market of software solutions.
- Provide primary operational support and engineering for multiple large distributed applications.
- Gather and analyze operating-system and application metrics for performance tuning and fault finding.
- Partner with development teams on testing and release procedures.
- Participate in system design consulting, platform management, and capacity planning.
- Create sustainable systems and services through automation and uplifts.
- Define and balance service-level objectives, feature delivery speed, and reliability.
- Drive incident response, resolution, and communication during outages and other critical incidents, including blameless postmortems.
Requirements
- 6–8 years of experience in a similar role focused on systems engineering, automation, and reliability.
- 4+ years of programming or scripting experience with Go, Python, .NET/C#, or Node.
- Proficiency in at least one programming language such as Python, Go, Java, or C#, plus scripting experience with Bash or PowerShell.
- Bachelor’s degree in computer science, engineering, or a related field, or equivalent experience.
- Deep understanding of cloud computing platforms such as AWS and services including EC2, ECS, Lambda, and DynamoDB.
- Experience with infrastructure-as-code tools such as CloudFormation and Terraform.
- Experience with CI/CD tools such as Jenkins, GitLab CI/CD, or CircleCI.
- Strong knowledge of Docker, Kubernetes, and microservices architecture.
- Experience with monitoring and observability tools such as Prometheus, Grafana, ELK Stack, and CloudWatch.
- Strong troubleshooting and problem-solving skills in distributed systems.
- Experience with incident management, blameless postmortems, and cross-functional incident response.
- Kubernetes certification, Grafana, AWS, Azure, and DevOps experience are advantageous.
Benefits
- NiCE-FLEX hybrid work model with two days in the office and three days working remotely each week.
- Collaborative, creative, fast-paced environment with learning, growth, and internal career opportunities.
- Equal opportunity employment across multiple roles, disciplines, domains, and locations.
Tech Stack
Categories
Site Reliability
About NICE
NiCE is transforming the world with AI that puts people first. Our purpose-built AI-powered platforms automate engagements into proactive, safe, intelligent actions, empowering individuals and organizations to innovate and act, from interaction to resolution. Trusted by organizations throughout 150+ countries worldwide, NiCE’s platforms are widely adopted across industries connecting people, systems, and workflows to work smarter at scale, elevating performance across the organization, delivering proven measurable outcomes.