1 day ago
Base Salary
$141k - $173k/yr
Responsibilities
- Design, deploy, scale, and maintain highly available Apache Airflow clusters for enterprise ETL/ELT workflows.
- Manage Kubernetes EKS/AKS clusters, Docker containers, and AWS/Azure cloud infrastructure.
- Drive infrastructure-as-code, GitOps, CI/CD automation, and self-hosted GitHub runner management.
- Build monitoring, logging, alerting, and telemetry systems using Grafana, Prometheus, Loki, and centralized log-management solutions.
- Architect and operate data infrastructure supporting Spark, EMR, Snowflake, Synapse, Kafka, and dbt.
- Deploy and maintain data discovery, metadata, lineage, schema-management, and governance tooling.
- Optimize cloud and data-platform compute and storage costs across AWS and Azure.
- Build internal tools, CLI utilities, and workflow templates for data and analytics engineers.
- Own platform modules from architecture and deployment through production operations and incident management.
- Define engineering standards for code quality, testing, security, access controls, IAM, and dynamic schema management.
- Develop the data-platform roadmap with technical, product, and business stakeholders.
- Mentor junior and mid-level data platform engineers as a subject matter expert.
Requirements
- 6+ years of hands-on professional experience in Data Platform Engineering, DevOps, or Site Reliability Engineering supporting big-data environments.
- Deep production expertise managing, tuning, scaling, and troubleshooting Apache Airflow, including Celery/Kubernetes Executors and DAG performance.
- Expert-level experience with Kubernetes, Helm, Terraform, Docker, ArgoCD, and GitHub Actions, including custom runner configurations.
- Production experience configuring alerting, metrics collection, and log aggregation with Grafana, Prometheus, and Loki.
- Operational and configuration experience with Spark, AWS EMR, Snowflake, Azure Synapse, dbt, and Apache Kafka.
- Hands-on experience with AWS and/or Azure, including demonstrated cloud-cost optimization work.
- Experience deploying or managing DataHub, Amundsen, or comparable metadata-management platforms.
- Strong programming or scripting skills in Python, Bash, Go, or SQL.
- Ability to take complex architectural requirements from concept through production-grade deployment.
- Commitment to automation, automated testing, CI/CD checks, and resilient infrastructure.
Benefits
- Remote work is available, but candidates must reside within 30 miles of Portland, Boston, Chicago, Dallas, the San Francisco Bay Area, or Seattle/Washington.
- Benefits include health, dental, and vision insurance, retirement savings, paid time off, health savings and flexible spending accounts, life and disability insurance, tuition reimbursement, and other benefits.
- Non-sales roles are typically eligible for a quarterly or annual bonus under the applicable plan.
Tech Stack
Apache AirflowApache KafkaApache SparkAWSAzureBashdbtDockerGitHub ActionsGoGrafanaHelmKubernetesPrometheusPythonSnowflakeSQLTerraform
Categories
Data EngineeringDevOps
About WEX
WEX is a public financial technology company that provides corporate payment solutions to fleets, healthcare/benefits administrators, and travel businesses. Its products include fuel and mobility cards, virtual card payments, and platforms for HSAs/FSAs/HRAs, with revenue from payment processing and software/service fees. Founded in 1983 and headquartered in Portland, Maine, WEX trades on the NYSE under WEX and was formerly known as Wright Express.
