
Staff Network Reliability Engineer, Service Assurance AIOps & Automation
Skylo Technologies2 months ago
Bengaluru, IndiaStaff+
Responsibilities
- Own the centralized network automation framework, reference architecture, multi-year roadmap, extensibility model, API contracts, schemas, governance, and operating model.
- Architect reusable health checks, diagnostic libraries, remediation actions, notification handlers, and governed workflow components for Network Operations and Network Engineering.
- Define DAG-based orchestration and closed-loop automation for diagnostics, root-cause classification, remediation, validation, evidence capture, and automated closure.
- Establish progressive automation tiers with confidence scoring, human approvals, promotion and demotion criteria, rollback paths, and auditability.
- Set standards for AI-assisted development and architect AI-powered diagnostic agents and alert-correlation systems.
- Architect real-time operational data pipelines, the data lake, DataOps controls, data-quality governance, and Grafana dashboards.
- Define production observability, CI/CD, testing, security scanning, identity, secrets management, audit logging, and staged rollout standards.
- Lead automation intake and prioritization, mentor automation developers, and measure coverage, MTTA, MTTR, false positives, overrides, and adoption.
Requirements
- 14+ years in platform automation, observability/SRE tooling, service assurance development, or network automation engineering, including 4+ years at a Principal, Staff-plus, or Architect level owning organization-wide automation or platform architecture.
- Demonstrated experience architecting centralized automation or platform frameworks consumed by multiple operations or engineering teams, including reference architecture, API and schema governance, and extensibility models.
- Expert production-grade Python development with modular architecture, testing, CI/CD integration, and code review; Go is a strong plus.
- Deep experience with DAG-based workflow orchestration such as Apache Airflow, Prefect, Dagster, or equivalent.
- Architecture-level data engineering experience with real-time streaming, data lakes, schema design, DataOps, and data-quality governance.
- Experience owning observability using Prometheus, PromQL, Grafana, VictoriaMetrics, and Loki or ELK.
- Experience setting standards for safe, auditable AI-assisted development using tools such as Claude Code, Claude Cowork, or GitHub Copilot.
- Experience with Kubernetes, GKE or EKS, Helm, ArgoCD, GitLab CI or Jenkins, security scanning, and rollback.
- Experience with alert correlation, event processing at scale, governance, audit, and security for production-changing automation.
- Ability to communicate architecture, operating models, roadmap, and automation impact to engineering leadership.
- Preferred experience includes telecom or NTN operations, 5G Core, RAN, agentic AI, ServiceNow or Jira automation, advanced Grafana development, chaos engineering, MLOps, OSS/BSS integration, and relevant certifications such as CKA, AWS/GCP Professional, Confluent Kafka, or Apache Airflow certification.
Benefits
- Competitive compensation packages including a stock option-based equity program.
- Comprehensive benefits plans.
- Monthly wellness and education reimbursement allowances.
- Generous time off policy, holidays, and the opportunity to temporarily work abroad.
- Flexible work approach with employees across three continents.
- Opportunity to develop and operate a commercial live direct-to-device satellite network.
- Access to a global team spanning software, hardware, chipsets, telecom, satellite, and network virtualization.
- Open, transparent, inclusive culture.
Tech Stack
Apache AirflowApache KafkaAWSGitLab CI/CDGoGoogle BigQueryGoogle Cloud PlatformGrafanaHelmJenkinsKubernetesPrometheusPython
Categories
Data EngineeringDevOps