7 months ago
Remote, United Kingdom or Cheltenham, United KingdomMid Level
Responsibilities
- Engineer distributed ingestion services that collect data from diverse sources and deliver structured outputs to downstream products.
- Build scalable, high-throughput batch and near-real-time processing components with attention to performance and cost.
- Design and evolve schemas, validation rules, versioning, and backward-compatible data contracts.
- Write maintainable code, unit and integration tests, and observability using metrics, logs, and tracing.
- Harden pipelines against failures, retries, rate limits, data drift, and infrastructure issues while improving reliability tooling and guardrails.
- Contribute to CI/CD, developer experience, design reviews, code reviews, incident retrospectives, and iterative delivery.
Requirements
- At least 2 years of experience building and operating production software systems.
- Fluency in at least one programming language, with Python or Node.js experience preferred.
- Experience debugging moderately complex systems and improving reliability and performance.
- Strong fundamentals in data structures, testing, version control, and Linux basics.
- Spark or PySpark, Hadoop ecosystem, workflow orchestration, search/indexing, Kubernetes, and infrastructure-as-code experience are desirable.
- A degree in Computer Science or a numerical discipline is desirable.
Benefits
- Remote working across the UK, with optional access to the Cheltenham office.
- 25 days of annual leave plus a birthday day off and bank holidays, rising to 30 days after five years.
- Private family healthcare, employee assistance programme, pension contributions, pension salary sacrifice, and enhanced maternity and paternity pay.
- 35-hour working week and access to company-provided technology including a MacBook Pro.
Tech Stack
AnsibleApache AirflowApache HBaseApache SparkArgo CDGitHub ActionsHelmKubernetesLinuxMongoDBNode.jsPythonTerraform
Categories
Data Engineering
