Site Reliability Engineer
Intermedia Intelligent Communications26 days ago
Remote, PortugalMid Level / Senior
Responsibilities
- Run and improve production environments supporting AI workloads, data pipelines, analytics applications, and customer-facing services.
- Build software and automation for cloud infrastructure, data platforms, model-serving infrastructure, and application services.
- Define and measure service-level indicators, service-level objectives, and error budgets for AI and analytics services.
- Implement end-to-end observability across metrics, logs, traces, data quality, AI-service performance, and customer impact.
- Monitor and optimize reliability, performance, capacity, and cost for batch, streaming, analytics, and inference workloads.
- Partner with data engineering and machine-learning teams to make ingestion, transformation, feature, training, deployment, and reporting workflows production-ready.
- Automate production-readiness checks and release processes for data pipelines, model and prompt releases, schema changes, analytics applications, and dashboards.
- Detect and resolve data-quality incidents involving stale or anomalous data, schema drift, broken lineage, and dependency failures.
- Design graceful degradation, dependency isolation, retry, fallback, and recovery procedures for AI, data, and downstream services.
- Improve the reliability and operational monitoring of AI-powered Voice and Unified Communications capabilities.
- Plan capacity and run performance, load, and resilience tests across AI, distributed data processing, and analytics services.
- Lead incident response and post-incident improvement efforts while reducing recurrence and recovery time.
- Improve engineering productivity through platform tooling, runbooks, self-service automation, and operational standards.
Requirements
- Bachelor’s degree in computer science, data engineering, software engineering, or another technical or scientific discipline, or equivalent practical experience.
- 4–7 years of experience in production operations, systems engineering, SRE or DevOps, software deployment, and distributed production systems.
- Experience with cloud infrastructure, containers, Kubernetes, distributed systems, and scalable compute and storage services.
- Experience operating data processing, orchestration, storage, or analytics technologies such as Kafka, Spark, Airflow, dbt, data warehouses, or comparable cloud services.
- Ability to use metrics, logs, traces, data-quality checks, freshness indicators, lineage, and service-level indicators to diagnose complex production issues.
- Experience with CI/CD, DataOps or MLOps practices, automated testing, controlled rollout, and rollback of data and AI service changes.
- Strong troubleshooting skills across Linux, applications, networks, APIs, data pipelines, and distributed service dependencies.
- Strong analytical problem-solving and cross-functional communication skills, with a proactive approach to reliability, performance, and continuous improvement.
Benefits
- Primarily remote work for candidates located in Portugal.
- Occasional visits to the office in Coimbra or Aveiro are required.
- The company plans to open an office in Porto in the future.
- The role offers opportunities to work with AI, analytics, cloud technology, and customer-facing communications platforms.
Tech Stack
Categories
Site Reliability
About Intermedia Intelligent Communications
Intermedia Intelligent Communications provides cloud-based unified communications and collaboration for businesses, spanning voice, video conferencing, chat/SMS, contact center, business email, file sharing, and backup. Its subscription UCaaS suite, including Intermedia Unite, is sold directly and through 7,500+ channel partners, serving over 150,000 businesses. Headquartered in Sunnyvale, California, Intermedia is privately held and backed by Madison Dearborn Partners.