6 days ago
Remote, Colombia or Bogotá, ColombiaMid Level
Responsibilities
- Design and develop scalable data processing solutions using Spark and Amazon EMR or comparable cloud-based platforms.
- Build and maintain batch and distributed data pipelines.
- Develop software for data transformation, feature preparation, and AI or machine learning workflow integration.
- Collaborate with engineering, AI, and product teams to operationalize data-driven and model-enabled use cases.
- Optimize pipeline performance, cost efficiency, scalability, and production reliability.
- Troubleshoot data and application issues across development and production environments.
- Contribute to architecture discussions, technical documentation, and engineering standards.
- Ensure solutions meet data quality, governance, and security expectations.
Requirements
- Require 4+ years of software engineering or data engineering experience.
- Require strong experience with Spark and distributed data processing.
- Require experience with Amazon EMR or a comparable cloud-based distributed data processing platform.
- Require proficiency in Java, Python, or a related programming language.
- Require exposure to AI or machine learning workflows, model integration, or data preparation for intelligent systems.
- Require understanding of scalable data architecture, performance optimization, debugging, collaboration, and practical architecture judgment.
- Prefer experience with Kafka, Airflow, data lakes, data warehouse ecosystems, MLOps, feature stores, AWS-native services, observability tooling, and enterprise environments.
Benefits
- Remote contract role for candidates located in LATAM, excluding Mexico.
- Full-time allocation of approximately 40 hours per week.
- U.S. Central Time coverage is required.
- Current contract end date is March 31, 2027.
Tech Stack
Categories
Data Engineering
