DataHub

Senior Software Engineer or Staff Engineer- Ingestion Framework

DataHub
Apply
3 hours ago
Remote, WorldwideSenior / Staff+

Responsibilities

  • Build and evolve the Python-based metadata-ingestion framework.
  • Design and ship connectors for modern data and AI platforms including Snowflake, Redshift, Kafka, and Databricks.
  • Improve ingestion scalability, resilience, observability, and cloud-native deployment across customer environments.
  • Apply AI to metadata inference, automated enrichment, and intelligent ingestion workflows.
  • Solve distributed-systems challenges involving large-scale metadata extraction, incremental updates, reliability, and performance.
  • Lead technical direction for a small engineering pod and mentor junior engineers.
  • Collaborate with product, platform, and customer-facing teams to turn data-integration needs into product capabilities.

Requirements

  • 8+ years of experience designing, building, and shipping production distributed backend systems.
  • Strong hands-on Python expertise across feature design, implementation, testing, deployment, and iteration.
  • Strong computer-science fundamentals and a degree in Computer Science or equivalent practical experience.
  • Deep understanding of API design, algorithmic complexity, reliability, and distributed-systems concepts.
  • Experience with modern data platforms such as Snowflake, Databricks, Kafka, or similar systems is highly valued.
  • Experience operating across multiple cloud environments is a strong plus.
  • Previous experience leading or mentoring a small team or engineering pod is a strong plus.
  • Ownership-oriented approach and ability to work effectively in a fast-moving, ambiguous environment.

Benefits

  • Remote work unless otherwise specified, with location flexibility for a home office or coworking space.
  • Monthly coworking stipend.
  • Equity for every team member.
  • Medical, dental, and vision coverage with premiums covered at 99% for employees and 65% for dependents.
  • Flexible Spending Accounts and optional Dependent Care FSA.
  • Carrot Fertility benefits and family-forming support for all U.S. employees.
  • Unlimited PTO and sick leave.
  • Role location is Toronto, Canada.

Tech Stack

Amazon RedshiftApache KafkaDatabricksPythonSnowflake

Categories

BackendData Engineering
DataHub

About DataHub

51-200 employees

DataHub transforms enterprise data into trusted context, enabling intelligent decision making by humans and AI agents. As the leading context management platform built on a thriving open-source community of 15,000+ members and adopted by thousands of organizations worldwide, DataHub Cloud delivers AI-powered discovery, governance, and observability in a unified platform, ensuring that context is always relevant, reliable, and continuously refreshed across the entire data estate. DataHub is backed by Bessemer Venture Partners, LinkedIn, and 8VC. Our key differentiators: * Scalability: DataHub offers best-in-class enterprise-grade scalability in connecting to over 100 data sources, offering an embeddable connector framework, and ingesting large volumes and high velocity of metadata. * Extensibility: DataHub’s highly extensible metadata model offers easy flexibility in adapting to an organization’s unique data landscape, entities, relationships, ownership, and custom metadata descriptors. * Completeness: DataHub Cloud’s unified platform uses AI-based enhancements and automations for discovery & understanding, quality management, and collaborative governance, allowing users and AI Agents to confidently use and manage data and AI assets. * Open-Source Advantage: Customers of DataHub benefit from the joint innovation, peer support, and growing skill base of an energized community of over 15,000 DataHub practitioners. The managed service, DataHub Cloud, offers dedicated support, improved performance and availability, and secure deployment options to ease adoption across an enterprise. For engineering teams deploying AI in production, DataHub delivers unified context infrastructure across all AI & data assets with enterprise-grade performance.