Member of Technical Staff, Data
The Inception Company6 months ago
San Mateo, CA, USAMid Level
Responsibilities
- Develop training data mixes using open-source datasets, synthetic data, and curated human feedback.
- Design and implement petabyte-scale data processing pipelines.
- Build web crawling, data ingestion, and real-time data processing systems for model training.
- Develop tools for distributed data storage, retrieval, and versioning.
- Create frameworks to evaluate data diversity, quality, and representativeness.
- Ensure data collection complies with privacy regulations.
- Support human annotation workflows and quality control processes.
Requirements
- BS, MS, or PhD in Computer Science, Machine Learning, or a related field, or equivalent experience.
- At least 3 years of experience building data processing pipelines at scale, particularly for AI/ML applications.
- Strong Python proficiency and experience with Apache Spark, Apache Beam, and Airflow.
- Familiarity with synthetic data generation and data augmentation techniques.
- Familiarity with web scraping, crawling technologies, and Common Crawl datasets.
- Understanding of machine learning fundamentals and experience with PyTorch or TensorFlow.
- Experience with SQL and NoSQL databases for structured and unstructured data.
- Preferred experience with large language models, tokenization, embeddings, model architectures, human annotation workflows, vector databases, embedding-based retrieval, distributed computing, large-scale storage, data privacy, and ethical AI practices.
Tech Stack
Categories
Data Engineering