The Inception Company

Member of Technical Staff, Data

The Inception Company
Apply
6 months ago
San Mateo, CA, USAMid Level

Responsibilities

  • Develop training data mixes using open-source datasets, synthetic data, and curated human feedback.
  • Design and implement petabyte-scale data processing pipelines.
  • Build web crawling, data ingestion, and real-time data processing systems for model training.
  • Develop tools for distributed data storage, retrieval, and versioning.
  • Create frameworks to evaluate data diversity, quality, and representativeness.
  • Ensure data collection complies with privacy regulations.
  • Support human annotation workflows and quality control processes.

Requirements

  • BS, MS, or PhD in Computer Science, Machine Learning, or a related field, or equivalent experience.
  • At least 3 years of experience building data processing pipelines at scale, particularly for AI/ML applications.
  • Strong Python proficiency and experience with Apache Spark, Apache Beam, and Airflow.
  • Familiarity with synthetic data generation and data augmentation techniques.
  • Familiarity with web scraping, crawling technologies, and Common Crawl datasets.
  • Understanding of machine learning fundamentals and experience with PyTorch or TensorFlow.
  • Experience with SQL and NoSQL databases for structured and unstructured data.
  • Preferred experience with large language models, tokenization, embeddings, model architectures, human annotation workflows, vector databases, embedding-based retrieval, distributed computing, large-scale storage, data privacy, and ethical AI practices.

Tech Stack

Apache AirflowApache BeamApache SparkGoogle BigQueryPythonPyTorchSQLTensorFlow

Categories

Data Engineering
The Inception Company

About The Inception Company

51-200 employees
Contact me