
Lead Data Engineer - Enterprise Data & Analytics - Remote
Mayo Clinic1 month ago
Rochester, MN, USAStaff+
Responsibilities
- Design, prototype, develop, review, and optimize production-grade data pipelines, data products, and platform capabilities.
- Lead data architecture, infrastructure redesign, process automation, data delivery optimization, and root-cause analysis.
- Establish architectural principles and design patterns while driving scalable, resilient, secure, and maintainable solutions.
- Provide hands-on technical leadership, participate in the Technical Review Board, and facilitate communication across engineering disciplines.
- Mentor and coach engineers, lead technical discussions, and investigate design approaches and technology prototypes.
- Collaborate with software engineers to analyze, develop, test, and deliver functional requirements.
- Build resilient data products, test automation suites, unit-test coverage, data-quality systems, monitoring, and observability.
- Lead data engineering teams using CI/CD practices and contribute significantly to production codebases.
Requirements
- Bachelor ’s degree in Computer Science, Engineering, or a related field with 6 years of experience, or an associate degree in a related field with 8 years of experience.
- At least 5 years of experience in data engineering, data science, or analytical modeling and at least 5 years using relational and NoSQL databases.
- Advanced proficiency in Python and SQL with experience building and supporting production-grade solutions.
- Experience with cloud platforms, cloud-agnostic architecture, modern data platform design, and enterprise-scale analytics, AI/ML, and operational workloads.
- Experience with Apache Spark, Hive, Airflow, Kafka, GCP Dataflow, Google BigQuery, FHIR APIs, Vertex AI, and Google Composer.
- Experience with distributed computing technologies such as Spark, Flink, Ray, or comparable frameworks.
- Experience with open data architecture technologies including Apache Iceberg, Delta Lake, and Apache Hudi, plus Parquet, Avro, and ORC formats.
- Experience implementing CI/CD, automated testing, infrastructure as code, observability, scalability, reliability, security, resiliency, and maintainability practices.
- Experience with Jenkins, GitHub Actions, Azure Pipelines, Jira, GitHub, SharePoint, and Azure Boards.
- Knowledge of Kubernetes and Docker in cloud environments and professional software engineering practices across the software development life cycle.
Tech Stack
Apache AirflowApache FlinkApache HiveApache KafkaApache SparkAWSAzureDockerGitHub ActionsGoogle BigQueryGoogle Cloud PlatformJenkinsKubernetesPythonSQL
Categories
Data Engineering