
Senior Engineer, Platform & Site Reliability
Intercontinental Exchange1 month ago
Responsibilities
- Provision, upgrade, and operate Amazon EKS Kubernetes clusters across AWS regions and environments.
- Manage cloud infrastructure declaratively with Crossplane and GitOps practices using ArgoCD.
- Operate the Istio service mesh, Envoy, ingress, canary rollouts, and mutual TLS traffic management.
- Design and operate observability solutions with Prometheus, Grafana, Jaeger, OpenTelemetry, Kiali, and Fluent Bit.
- Improve reliability through capacity planning, autoscaling, incident response, on-call practices, chaos engineering, SLOs, alerting, and dashboards.
- Administer platform services including cert-manager, external-dns, external-secrets, sealed-secrets, and AWS Load Balancer Controller.
- Build and maintain Azure DevOps CI/CD pipelines and automated security scanning such as SonarQube.
- Develop Java Spring platform tooling and microservices plus React TypeScript micro frontends.
- Design APIs and automation supporting platform capabilities and self-service for product teams.
- Troubleshoot test and production failures, lead root-cause analysis, and develop or review automated reliability tests.
- Write technical specifications and operational runbooks and participate in reliability and architecture design ceremonies.
- Mentor or guide less experienced site reliability and software engineers.
Requirements
- Bachelor's degree or equivalent combination of education, training, or work experience.
- 5+ years of site reliability, platform, DevOps, or software engineering experience.
- Hands-on production experience operating Kubernetes and cloud-native technologies, preferably Amazon EKS on AWS.
- Experience with infrastructure as code, GitOps practices, and CI/CD pipeline development and maintenance.
- Working proficiency in Java or TypeScript/JavaScript for automation and platform tooling.
- Experience with observability frameworks, SLO/SLI definition, alerting, service meshes, secrets management, certificate automation, chaos engineering, resilience testing, and RESTful services is preferred.
- Experience with Java JVM workloads, microservices, Postgres SQL databases, PL/SQL, React, and source-code tools such as Azure DevOps, TFS, Jira, or Git is preferred.
- Familiarity with Test-Driven Development, Behavior-Driven Development, Unit Tests, Component Tests, Scenario Tests, and Agile SDLC practices is preferred.
Tech Stack
AmbassadorAWSGitGrafanaIstioJavaJavaScriptKubernetesOpenShiftPostgreSQLPrometheusReactSonarQubeTensorFlowTypeScript
Categories
DevOpsSite Reliability
About Intercontinental Exchange
ICE (NYSE: ICE) connects people to data, technology and expertise that create opportunity and inspire innovation. For terms of use, visit www.ice.com/privacy-security-center/terms-of-use