Site Reliability Engineer
Intermedia Intelligent Communications1 day ago
Remote, PortugalMid Level / Senior
Responsibilities
- Run and improve production environments by monitoring availability and evaluating overall system health.
- Build software, automation, and systems to manage platform infrastructure and applications.
- Improve reliability, quality, scalability, and time-to-market across cloud software solutions.
- Measure and optimize system performance, capacity, failure modes, and bottlenecks.
- Provide operational support and engineering for large distributed applications and services.
- Analyze operating system, application, network, and service metrics for performance tuning and fault isolation.
- Partner with development teams on testing, release procedures, and production-readiness practices.
- Participate in system design, platform management, capacity planning, incident response, and post-incident improvement.
- Automate infrastructure improvements and reduce operational toil.
- Define and use service-level indicators, service-level objectives, and error budgets to balance feature delivery and reliability.
- Improve reliability of Voice/UC platforms and integrations, including real-time signaling, media flows, call quality, latency, jitter, packet loss, failover, and service availability.
- Operationalize AI-enabled communications capabilities and build observability across Voice/UC and AI service paths.
- Design and test graceful degradation, dependency isolation, retry and fallback patterns, and recovery procedures.
- Automate validation and production-readiness checks for Voice/UC and AI integrations.
Requirements
- Bachelor's degree in computer science or another highly technical or scientific discipline, or equivalent practical experience.
- 4–7 years of experience in production operations, systems engineering, SRE/DevOps, CI/CD implementation, software deployment, and production-system maintenance.
- Experience with Agile methodologies, DevOps practices, CI/CD pipelines, infrastructure automation, and production monitoring and observability.
- Experience with distributed systems, cloud infrastructure, containers, and dynamic resource management frameworks such as Kubernetes.
- Experience with distributed storage technologies such as NFS, HDFS, or S3, or comparable cloud storage technologies.
- Hands-on troubleshooting experience across Linux, applications, networks, APIs, and distributed service dependencies.
- Working knowledge of Voice/UC and real-time communications concepts such as SIP, RTP/SRTP, WebRTC, SBCs, and media services.
- Experience supporting or integrating AI-enabled services, APIs, or workflows; familiarity with speech/voice AI, machine-learning services, or LLM-based applications is preferred.
- Ability to use metrics, logs, traces, and service-level indicators to diagnose production issues and drive reliability improvements.
- Proactive problem-solving approach with strong cross-functional communication skills.
Tech Stack
Categories
Site Reliability
About Intermedia Intelligent Communications
Intermedia Intelligent Communications provides cloud-based unified communications and collaboration for businesses, spanning voice, video conferencing, chat/SMS, contact center, business email, file sharing, and backup. Its subscription UCaaS suite, including Intermedia Unite, is sold directly and through 7,500+ channel partners, serving over 150,000 businesses. Headquartered in Sunnyvale, California, Intermedia is privately held and backed by Madison Dearborn Partners.