4 hours ago
Remote, Peru +6 moreSenior
Responsibilities
- Collaborate with development teams to implement application monitoring, alerting, dashboards, and APM instrumentation.
- Lead the implementation, configuration, and optimization of APM solutions.
- Apply Azure Monitor, Application Insights, New Relic, and Log Analytics/KQL observability practices.
- Implement code-level instrumentation, distributed tracing, and structured logging.
- Design and maintain application monitoring dashboards and operational health metrics.
- Define SLIs, SLOs, and alerting strategies based on latency, error rates, traffic, and resource saturation.
- Improve monitoring and alerting through production insights and incident learnings.
- Participate in production-readiness reviews and identify operational risks, observability gaps, and failure scenarios.
- Support incident analysis and post-incident improvements through enhanced telemetry and monitoring.
- Partner with engineering teams to ensure applications are reliable, scalable, and production-ready.
Requirements
- 5+ years of experience in Site Reliability Engineering, Cloud Operations, or related roles.
- At least 1 year of hands-on experience with dbt, Databricks, and SQL.
- Strong experience supporting and operating applications in Microsoft Azure IaaS environments.
- Hands-on experience with application observability, monitoring, and reliability engineering practices.
- Experience implementing and managing APM solutions such as Application Insights or New Relic.
- Experience designing dashboards and monitoring solutions with Azure Monitor, Application Insights, and Log Analytics/KQL.
- Familiarity with CI/CD environments including Azure DevOps and GitHub Actions.
- Understanding of cloud-native architectures and distributed application systems.
- Experience with incident analysis, root-cause investigation, and proactive problem prevention.
- Strong verbal and written English communication skills for collaboration with global teams.
- Preferred experience with PowerShell and/or Bash scripting and automation.
- Preferred knowledge of scalability, availability, and resilience patterns in modern cloud environments.
- Preferred experience driving production-readiness and operational-excellence initiatives.
- Preferred exposure to reliability engineering practices in enterprise-scale environments.
Benefits
- Remote work arrangement.
- Position is available to legal residents of Peru, Colombia, Bolivia, Costa Rica, Mexico, and Brazil.
About Encora
Encora provides software and digital product engineering services for technology companies and enterprises, including cloud-native development, data engineering, and QA. It operates a global delivery model and supports outsourced product development and managed engineering teams. Headquartered in Santa Clara, California, Encora is privately held and was acquired by Coforge; it previously raised $200 million in private equity funding in 2019.
