
Data Engineer I (R-19886)
Dun & Bradstreet1 day ago
Hyderābād, IndiaEntry Level
Responsibilities
- Write, review, test, and maintain SQL and Python code for data tools, automations, and ingestion and transformation workflows.
- Automate manual processes and develop reusable data-processing components, web-scraping solutions, and API integrations.
- Build and support scalable ETL/ELT pipelines for structured and unstructured data.
- Evaluate, validate, document, and monitor data sources for research, managed-services, and AI-enabled workflows.
- Implement AI-enabled workflows using orchestration frameworks, large language models, embeddings, retrieval-augmented generation, vector databases, and prompt engineering.
- Perform monitoring, exception handling, data maintenance, database administration, performance tuning, and production issue resolution.
- Validate, reconcile, clean, and maintain data integrity while following data governance, security, and operational standards.
- Develop documentation including data dictionaries, source assessments, data-flow diagrams, data mappings, runbooks, and data lineage.
- Collaborate with cross-functional technical and non-technical stakeholders and conduct knowledge-exchange sessions.
Requirements
- Strong SQL and Python skills with demonstrated daily coding experience.
- Experience with Playwright, Selenium, web-data collection, web scraping, and data-ingestion, transformation, and ETL/ELT workflows.
- Ability to collect and interpret data from multiple sources, including GCS and S3, and formats including delimited files, XML, JSON, and PDF.
- Working knowledge of data systems and databases used to maintain data pipelines, including BigQuery and cloud platforms such as AWS and GCP.
- Experience with Power BI, Tableau, or other dashboard tools.
- Experience managing stakeholders and project plans, with proficiency in Microsoft Office Suite.
- Hands-on experience implementing AI solutions with LangChain or an equivalent orchestration framework.
- Exposure to large language models, prompt engineering, retrieval-augmented generation, embeddings, vector databases, AI agents, or graph databases.
- Knowledge of Data Operations methodologies, data management approaches, ServiceNow, and/or Jira.
- Experience with NoSQL technologies, SQL Server administration, R programming, web technologies, and data mapping from multiple sources.
- Willingness and demonstrated ability to learn new technologies and platforms.
Tech Stack
Categories
Data Engineering