1 day ago
Chennai, IndiaStaff+

Responsibilities

  • Develop and maintain robust data pipelines for structured and unstructured data used by data scientists to build machine learning models.
  • Build and maintain data engineering tools, shared utility libraries, and data quality libraries used by multiple teams.
  • Work with Databricks, Delta Lake, workflows, job clusters, the Databricks CLI, and Databricks Workspace.
  • Design and develop solutions using cloud data warehouses and distributed data processing frameworks.
  • Develop and use REST APIs, Python-based API frameworks, containers, orchestration tools, CI/CD tools, and version control systems.
  • Apply software engineering practices throughout the design and development lifecycle and collaborate with cross-functional technology and business teams.

Requirements

  • Bachelor’s or master’s degree in Computer Science, Data Science, Engineering, or a related field.
  • 8+ years of experience in the stated technical or related fields.
  • 4+ years of experience with Python, SQL, PySpark, and Bash scripting.
  • 3+ years of experience developing and maintaining data pipelines for structured and unstructured data.
  • 3+ years of experience with cloud data warehousing platforms and distributed frameworks such as Spark.
  • 2+ years of hands-on experience with the Databricks platform for data engineering.
  • Detailed knowledge of Delta Lake, Databricks Workflow, Job Clusters, Databricks CLI, and Databricks Workspace.
  • Understanding of the machine learning lifecycle, data mining, and ETL techniques.
  • Familiarity with scikit-learn and XGBoost and with code used for model training and scoring.
  • Experience with REST APIs and Python API development frameworks such as Flask or FastAPI.
  • Experience with containerization frameworks such as Docker or Kubernetes.
  • Experience with CI/CD tools such as Jenkins, version control systems such as GitHub or Bitbucket, and orchestration tools such as Airflow or Prefect.
  • Strong understanding of software engineering principles and strong communication and collaboration skills.

Benefits

  • The position can be based in Chennai or Gurgaon.

Tech Stack

Amazon RedshiftApache AirflowApache SparkBashDatabricksDockerFastAPIFlaskJenkinsKubernetesPythonscikit-learnSnowflakeSQLXGBoost

Categories

Data EngineeringDevOps
The Guardian Life Insurance Company of America

About The Guardian Life Insurance Company of America

10,000+ employees

The Guardian Life Insurance Company of America is a U.S. mutual insurer that provides life and disability insurance, dental and vision plans, annuities, and workplace benefits to individuals, families, and employers. It earns revenue through premiums, investment income, and fee-based wealth and retirement solutions, distributed via financial professionals, employers, and digital channels. Founded in 1860 and headquartered in New York City, it operates nationwide and is owned by its policyholders.

Contact me