2 months ago
Remote, Poland +8 moreSenior
Responsibilities
- Own and standardize the Jira Service Management-based Tech Ops Incident Management Platform across production fleets.
- Automate incident-resolution workflows using AI tooling, including runbook generation and assignment.
- Design and document Tech Ops incident processes to support a clean handover.
- Build fleet-management and task-automation capabilities for global 24/7 remote operations.
- Own and customize the Grafana-based data-driven observability layer across internal technology teams.
- Independently drive architecture decisions with infrastructure and engineering leadership.
Requirements
- Senior-level experience as an AWS Cloud Engineer or Site Reliability Engineer.
- Strong hands-on experience with metric-driven observability, monitoring, and alerting at scale.
- Fluent hands-on Python experience for tooling and automation.
- Hands-on Terraform experience for Infrastructure as Code.
- Proven experience designing, architecting, and owning production systems end to end.
- Ability to work independently without detailed specifications or heavy direction.
- Experience with Jira Service Management or similar ITSM/incident platforms is a plus.
- Experience with Grafana dashboarding pipelines at scale is a plus.
- Exposure to AI-assisted operations tooling, AI-Ops, or runbook automation is a plus.
- Prior experience in hardware, robotics, or IoT fleet environments is a plus.
Benefits
- Remote workplace.
- Full-time freelance engagement for a 4–5 month project.