Senior Site Reliability Engineer - PSRE
Arcesium LLC3 months ago
Responsibilities
- Lead and coordinate incidents and critical platform issues during New York business hours.
- Monitor application and infrastructure health, analyze trends, identify risks, and implement preventative reliability measures.
- Troubleshoot complex issues across application, infrastructure, and network layers and identify root causes.
- Communicate incident updates and system status to engineering teams and stakeholders.
- Develop tools and scripts to automate tasks, improve operational efficiency, and reduce manual intervention.
- Contribute to the improvement of SRE practices, tools, processes, and team knowledge sharing.
Requirements
- Up to five years of experience in Site Reliability Engineering, DevOps, or Production Engineering.
- Incident management experience, including triaging, escalation, and resolution of high-severity outages.
- Proficiency in at least one coding language, specifically Python or Java, for automation and debugging.
- Hands-on experience with Kubernetes for managing and orchestrating containerized applications.
- Cloud experience, preferably AWS, including exposure to EC2, S3, Lambda, and CloudWatch.
- Strong troubleshooting, problem-solving, communication, prioritization, and multitasking skills.
- Fluency in spoken and written English.
- Legal right to work in the country.
- Terraform or CloudFormation experience is desirable.
- Experience with Datadog, Prometheus, or Grafana is desirable.
- Familiarity with web application architectures, CI/CD pipelines, and DevOps workflows is desirable.
Benefits
- Flexible hybrid work arrangement
- Casual dress code
- Continuous learning and development opportunities
- Collaborative and innovative work culture
- Competitive benefits package
- Modern office located at Avenida da Liberdade in Lisbon