Responsibilities
- Design, architect, operate, and modernize highly available, scalable, secure, and resilient cloud platform infrastructure.
- Define and implement reliability, observability, disaster recovery, capacity management, and operational excellence strategies.
- Design, deploy, and operate NATS messaging infrastructure and configure HAProxy for load balancing, routing, SSL termination, and high availability.
- Manage GCP services including GKE, Compute Engine, Cloud Storage, BigQuery, Pub/Sub, Cloud SQL, and Composer/Airflow.
- Implement infrastructure as code with Terraform and Terragrunt and automate provisioning, deployment, compliance, and operational workflows.
- Design and maintain enterprise CI/CD pipelines using Jenkins, GitLab CI, GitHub Actions, or equivalent tools.
- Implement monitoring, logging, tracing, and alerting with Prometheus, Grafana, ELK/OpenSearch, VictoriaMetrics/VictoriaLogs, and GCP Monitoring.
- Troubleshoot cloud, networking, OLT/ONT, and customer-facing application issues and lead root cause analysis for critical incidents.
- Develop technical roadmaps, contribute to design reviews and architecture governance, and mentor engineers.
- Collaborate with product, engineering, SRE, network operations, TAC, and customer success teams.
Requirements
- Bachelor’s degree in Computer Science, Engineering, Telecommunications, or equivalent practical experience.
- 8+ years of experience designing and operating large-scale distributed systems in production environments.
- 5+ years of hands-on experience with public cloud platforms, preferably Google Cloud Platform.
- Strong experience with Kubernetes and containerized platforms.
- Deep understanding of TCP/IP, routing and switching, NAT, DNS, DHCP, firewalls, VPN technologies, and load balancing.
- Hands-on experience configuring, tuning, troubleshooting, and deploying HAProxy in high-availability environments.
- Hands-on experience with NATS messaging systems and event-driven architectures.
- Experience supporting telecom and broadband infrastructure, including OLTs, ONTs, routers, and access network technologies.
- Experience implementing infrastructure as code using Terraform.
- Proficiency in Python, Go, Bash, or similar scripting or programming languages.
- Experience with CI/CD pipelines, DevOps platforms, monitoring, telemetry, observability frameworks, troubleshooting, and incident management.
- Preferred experience with broadband access networks, PON technologies, service provider environments, Kubernetes networking, service mesh technologies, Kafka, Redis, Elasticsearch/OpenSearch, distributed databases, SRE, platform engineering, or telecom cloud operations.
- Knowledge of cloud and telecom security best practices and Google Cloud Professional Certification(s) is preferred.
Benefits
- The position may be eligible for a bonus as part of the total compensation package.
- Benefits information is provided through the employer’s benefits program.
- The Canada base pay range is CAD 141,000–240,000 annually, with pay determined by location and other job-related factors.
Tech Stack
Categories
About Calix
Calix is an AI platform company that enables service providers to transform their operations and accelerate delivery of differentiated experiences—so they can compete and win in the markets and communities they serve. Through the AI-native Calix One platform, service providers can securely and privately activate agentic-AI alongside their human teams to acquire new subscribers, grow existing subscriber revenue, and build loyalty across residential, business, municipal, and MDU markets. More than 1,200 customers of all sizes leverage the Calix One platform, which has evolved over 15 years at an investment of more than $2 billion. Calix innovation cycles are underpinned by a strong financial balance sheet and a people‑first culture that routinely earns broad industry recognition—winning 81 culture and innovation awards since 2025 alone, as well as Fortune’s 100 Best Companies to Work For® in 2026.
