2 months ago
Base Salary
$141k - $217k/yr
Responsibilities
- Design and improve observability, monitoring, alerting, and incident response across the Unified Call platform.
- Build and mature deployment processes that improve reliability and reduce operational risk.
- Partner with Platform Engineering and Unified Call engineering to improve system resiliency and uptime.
- Analyze system architecture, identify operational risks, and implement preventive improvements.
- Develop dashboards, automation, and operational tooling for confident production operations.
- Help establish Site Reliability Engineering best practices as the organization scales.
Requirements
- At least 5 years of experience as a Site Reliability Engineer, Infrastructure Engineer, Platform Engineer, or in a similar role.
- Strong experience with observability platforms, with Datadog strongly preferred.
- Experience supporting Kubernetes-based cloud infrastructure, preferably AWS.
- Experience designing monitoring, alerting, deployments, and operational automation for production systems.
- Experience with message queues, event-driven architectures, or messaging platforms such as RabbitMQ or Kafka.
- Working knowledge of cloud networking fundamentals.
- Ability to build systems from the ground up, take ownership of ambiguous technical problems, lead operational improvements, and collaborate across engineering teams.
- Telephony or communications infrastructure experience is a strong bonus.
- Experience supporting highly available, mission-critical distributed systems is a bonus.
Benefits
- Hybrid role with an expectation of four days per week in the Manhattan office.
- Competitive salary and 401k with employer match.
- Discretionary time off.
- Paid parental leave for all employees.
- Medical, dental, and vision plans.
- Fitness programs, emotional and development programs, and office snacks.
Tech Stack
Categories
DevOpsSite Reliability
