4 hours ago
Responsibilities
- Design, build, and maintain scalable Python and SQL data and ML pipelines for batch and streaming workloads.
- Own Databricks platform infrastructure, including Delta Lake table architecture, Unity Catalog governance, Databricks Workflows orchestration, and compute optimization.
- Build and maintain end-to-end ML pipelines for feature engineering, model training, experiment tracking, and model deployment or serving.
- Collaborate with data scientists to operationalize models in production-grade ML systems.
- Define and enforce ingestion, data modeling, medallion architecture, and pipeline reliability standards.
- Implement data quality, observability, and monitoring frameworks for platform health and data trustworthiness.
- Optimize pipelines for performance, cost, and reliability at scale using Spark and PySpark.
- Evaluate, integrate, and govern platform tooling and data sources within the Databricks ecosystem.
- Contribute to architectural decisions and the long-term data platform roadmap.
- Participate in code reviews, technical design discussions, and engineering standards.
- Mentor junior engineers and improve platform and data engineering practices.
- Document platform architecture, pipeline design, and operational runbooks.
Requirements
- At least 6 years of experience in data engineering, data platform, or ML engineering roles.
- Strong proficiency in Python and SQL, with experience building production-grade data pipelines.
- Hands-on expertise with Databricks, Delta Lake, Unity Catalog, Databricks Workflows, PySpark, and the broader Databricks ecosystem.
- Experience building and maintaining production ML pipelines covering feature engineering, model training, experiment tracking, and model deployment.
- Familiarity with MLflow or comparable experiment tracking and model registry tools.
- Experience with cloud data platforms such as AWS, Azure, or GCP.
- Strong understanding of data modeling, dimensional design, and analytics-friendly data architecture.
- Experience with batch and incremental or CDC pipeline patterns.
- Proficiency with Git, version control, and CI/CD practices for data and ML workflows.
- Strong engineering judgment focused on reliability, maintainability, and cost.
- Clear communication and comfort working with technical and non-technical stakeholders.
- Preferred experience with streaming or near-real-time pipelines using Kafka, Kinesis, or Spark Structured Streaming.
- Preferred familiarity with Databricks Feature Store, Feast, or Tecton.
- Preferred experience with LLM pipelines, RAG architectures, or AI/BI tooling such as Genie and AI Functions.
- Preferred knowledge of data quality and observability tools such as Great Expectations and Monte Carlo.
- Preferred exposure to dbt or similar SQL-based transformation frameworks.
- Preferred infrastructure-as-code experience with Terraform or Databricks Asset Bundles.
- Prior experience mentoring engineers or shaping platform standards is preferred.
Benefits
- Generous time off policies, top-shelf benefits, and education, wellness, and lifestyle support.
Tech Stack
Categories
Data EngineeringML Engineering
About DigiCert
As the leading provider of digital trust, DigiCert solutions help you engage online with confidence, knowing your data and connections are secure. From enterprise solutions to software security, identity management to IoT, and DNS to quantum cryptography, DigiCert is trusted around the globe by individuals, businesses, governments, and 90% of Fortune 500 companies.
