2 months ago
Remote, SerbiaMid Level
Responsibilities
- Provide day-and-night on-duty service coverage and troubleshoot and resolve production incidents, involving vendors or third-party support when needed.
- Direct issues and queries to the appropriate departments and maintain detailed infrastructure records and Root Cause Analyses.
- Collaborate with teams to define monitoring needs and configure monitoring and observability systems accordingly.
- Design, adjust, and refine alerts and dashboards to improve relevance, clarity, and operational visibility.
- Build and maintain integrations between monitoring systems and platforms such as Jira and Opsgenie.
- Create and update a knowledge base covering configurations, alert processes, troubleshooting guidance, and user manuals.
- Identify opportunities to automate repetitive monitoring and support work, including with AI-assisted approaches where appropriate.
- Stay current with monitoring and observability trends and best practices.
Requirements
- At least 3 years of experience as a Systems Engineer, SRE, DevOps Engineer, or Monitoring Support Engineer at L2 level or higher.
- Good understanding of Linux-like, particularly Debian-based, operating systems.
- Experience with containerization, virtualization, and orchestration using LXC/LXD, Docker, and Kubernetes.
- Development experience in a scripting language such as Bash, Python, or Go, and familiarity with REST API.
- Knowledge of basic database concepts, including transactions and WAL; PostgreSQL experience is preferred.
- English proficiency at Intermediate (B1) level or higher and Russian proficiency at Upper-Intermediate (B2) level or higher.
- Practical interest in AI-assisted tools for troubleshooting, automation, documentation, and operational efficiency.
- Ability to critically evaluate and validate AI-generated output before production use, with an understanding of AI risks and limitations in infrastructure operations.
Benefits
- Private health insurance.
- Sports benefits and a comprehensive mental health program.
- Free online English lessons and local language courses.
- Paid time off and maternity leave support.
- Referral program rewards.
- Upskilling, internal workshops, professional conferences, and corporate events.
- Day-and-night on-duty shift coverage is part of the role.
Tech Stack
Categories
DevOpsSite Reliability
