IMC

Site Reliability Engineer - Data Platform

IMC
Apply
7 hours ago
Amsterdam, NetherlandsSenior

Responsibilities

  • Design, implement, and operate distributed data platforms and critical data services.
  • Improve observability to identify and resolve issues proactively.
  • Build automation that reduces operational toil and enables systems to scale.
  • Support and own the reliability of services including HDFS, Kafka, and Dremio.
  • Drive long-term architectural improvements across the data platform.
  • Deploy, configure, troubleshoot, and performance-tune systems across bare-metal Linux and Kubernetes.

Requirements

  • Strong experience managing distributed data platforms such as Kafka, Hadoop, Spark, and Dremio, including installation, debugging, and performance tuning.
  • Hands-on experience deploying, configuring, and orchestrating software on Linux and Kubernetes.
  • Strong infrastructure-as-code experience, with Ansible preferred.
  • Proficient programming experience in Python.
  • Ability to read, write, and tune SQL queries.
  • Ability to read Java source code and tune and debug running JVMs.
  • Familiarity with data lakehouse technologies such as Iceberg or Delta Lake and query engines such as Dremio, Presto, or Trino.
  • Exposure to workflow orchestration tools such as Airflow or Dagster.
  • Proactive approach to preventing operational issues and ability to work across teams with minimal oversight.

Tech Stack

AnsibleApache AirflowApache FlinkApache HadoopApache KafkaApache SparkBashClickHouseGrafanaHelmJavaKubernetesLinuxPrestoPrometheusPuppetPythonSQL

Categories

DevOpsSite Reliability
Contact me