Member of Technical Staff, Data Infrastructure
The Inception Company6 months ago
San Mateo, CA, USAMid Level
Responsibilities
- Design, build, and operate scalable, fault-tolerant infrastructure for distributed LLM training across compute, orchestration, and storage.
- Develop high-throughput data ingestion, processing, transformation, cataloging, deduplication, quality-checking, and search systems.
- Build web crawling, data ingestion, real-time processing, storage, retrieval, and versioning tools for model training operations.
- Partner with researchers to accelerate experiments, develop datasets, improve infrastructure efficiency, and generate insights from data assets.
- Ensure data collection complies with privacy regulations.
Requirements
- Bachelor’s, master’s, or doctoral degree in Computer Science, Machine Learning, or a related field, or equivalent experience.
- At least 3 years of experience building data processing pipelines at scale, particularly for AI or machine learning applications.
- Strong proficiency in Python and experience with Apache Spark, Apache Beam, and Apache Airflow.
- Familiarity with synthetic data generation, data augmentation, web scraping, crawling technologies, and Common Crawl datasets.
- Understanding of machine learning fundamentals and experience with PyTorch or TensorFlow.
- Experience using SQL and NoSQL databases for structured and unstructured data.
- Preferred: experience with large language models, tokenization, embeddings, model architectures, human annotation workflows, quality control, vector databases, embedding-based retrieval, distributed computing, and large-scale storage systems.
- Preferred: knowledge of data privacy regulations and ethical AI practices.
Tech Stack
Categories
Data Engineering