
Staff Databricks Engineer, Site Reliability Engineering
Kinaxis Inc.7 hours ago
Chennai, IndiaStaff+
Responsibilities
- Transition legacy systems to a modern cloud-native data stack using tools including Informatica, Airflow, Postgres, Snowflake, dbt, BigQuery, Looker, Databricks, Power BI, Datadog, and Grafana.
- Deploy, upgrade, monitor, troubleshoot, and support applications, services, operating systems, and production cloud infrastructure.
- Apply software engineering principles to automate operations, implement self-healing systems, and improve monitoring and service reliability.
- Deliver cloud reference architectures, operational plans, and scalable and secure production environments.
- Develop and manage cloud automation using orchestration capabilities, scripting languages, and APIs.
- Participate in an on-call rotation, investigate incidents, determine root causes, document resolutions, and coach junior employees.
- Design and manage reporting solutions supporting cloud infrastructure and optimize cloud resources for cost effectiveness.
- Evaluate emerging cloud technologies and lead critical escalations with Global Customer Care teams.
Requirements
- Bachelor’s or master’s degree in engineering with a computer science specialization or related discipline, or demonstrated equivalent experience.
- Prior experience in a site reliability engineering role.
- Deep hands-on experience with Snowflake, Databricks, Informatica, dbt, BigQuery, Looker, and/or Power BI.
- 10+ years of experience deploying and monitoring distributed systems, infrastructure, and cloud services across public cloud platforms such as GCP, Azure, or AWS.
- Familiarity with VMware ESXi is considered an asset.
- Experience with observability tools such as Logstash, Datadog, and Grafana.
- Strong knowledge of system design and operational and reliability trade-offs.
- Extensive scripting experience with Ansible, PowerShell, Bash, and Python.
- Practical experience building and managing Terraform infrastructure as code, Ansible configuration management, Git and GitOps CI/CD solutions, Argo CD, Docker, Kubernetes, Helm, Datadog, Prometheus, and ELK.
- Strong communication and documentation skills.
Benefits
- Flexible vacation and company-wide Kinaxis Days.
- Flexible work options.
- Physical and mental well-being programs, including regularly scheduled virtual fitness classes.
- Mentorship, training, and career development programs.
- Recognition programs, referral rewards, and hackathons.
- Inclusive recruitment process with accommodations available upon request.
Tech Stack
AnsibleApache AirflowArgo CDAWSAzureBashDatabricksDatadogdbtDockerGitGoogle BigQueryGoogle Cloud PlatformGrafanaHelmInformaticaKubernetesLogstashPostgreSQLPowerShellPrometheusPythonSnowflakeTerraform
Categories
Data EngineeringSite Reliability
About Kinaxis Inc.
Kinaxis builds RapidResponse, a cloud-based supply chain planning and orchestration platform used for S&OP, demand/supply planning, and risk management by enterprises in automotive, life sciences, high tech, and industrial sectors. It sells subscription software and related services to manage end-to-end supply chains. Founded in 1984 and headquartered in Ottawa, the company is publicly traded on the Toronto Stock Exchange (KXS).