
Gen AI Data Engineer
Tiger Analytics Inc.over 1 year ago
Remote, United StatesSenior
Responsibilities
- Architect and build distributed data systems and large-scale data warehouses capable of processing petabytes of data.
- Design and manage real-time and batch ingestion pipelines and data platforms.
- Build document-ingestion, indexing, retrieval, and vector-search pipelines for scalable RAG solutions.
- Develop and maintain data persistence and retrieval systems across relational, vector, bucket, graph, and knowledge-graph stores.
- Integrate external databases, APIs, and knowledge graphs into RAG systems.
- Create ontologies, schema-level constructs, and Open Cypher-based knowledge-retrieval solutions.
- Curate and collect data from structured, unstructured, traditional, and non-traditional sources.
- Conduct experiments, evaluate RAG workflow effectiveness, and iterate for improved performance.
- Monitor and optimize system performance, reliability, scalability, and security.
- Apply infrastructure automation, containerization, CI/CD, disaster recovery, and cloud security practices.
Requirements
- 8+ years of experience in data engineering, platform engineering, or related fields.
- Deep expertise designing and building distributed data systems and large-scale data warehouses.
- Proven experience architecting data platforms for petabyte-scale processing and real-time and batch ingestion.
- Strong experience building data pipelines for document ingestion, indexing, and retrieval for RAG solutions.
- Proficiency with Python, SQL, PySpark, Apache Airflow, cloud platforms, and GitHub.
- Experience with Snowflake, NoSQL, Neo4j, Hadoop, Spark, Streamlit, APIs, and vector databases.
- Experience with information retrieval, vector search, FAISS, Pinecone, Elasticsearch, or Milvus.
- Experience with graphs, graph algorithms, LLMs, optimization algorithms, relational databases, and diverse data formats.
- Experience building knowledge graphs and ontologies, including higher-level classes, punning, property inheritance, and Open Cypher.
- Experience with AWS services including SageMaker, Lambda, OpenSearch, S3, RDS, AWS Batch, and CloudFormation, or with GCP services including Vertex AI, BigQuery, and GKE.
- Advanced skills in Terraform and CloudFormation, plus experience with Docker, Kubernetes, Jenkins, and GitHub Actions.
- Experience with Airflow DAGs, AutoSys, CronJobs, unstructured data, monitoring, scalability, disaster recovery, and cloud security.
- Familiarity with generative AI tools and techniques and a basic understanding of machine learning concepts.
- Strong analytical, problem-solving, communication, presentation, and collaboration skills.
Benefits
- Remote, full-time position in the United States.
- Opportunity for significant career development in a fast-growing, challenging entrepreneurial environment with a high degree of individual responsibility.
Tech Stack
Apache AirflowApache HadoopApache SparkAWSDockerElasticsearchGitHub ActionsGoogle BigQueryGoogle Cloud PlatformJenkinsKubernetesLinuxNeo4jPythonSnowflakeSQLTerraform
Categories
Data Engineering