3 months ago
Bengaluru, IndiaSenior
Responsibilities
- Build highly scalable, available, and fault-tolerant batch and streaming data processing systems handling tens of terabytes of daily ingestion and petabyte-scale warehouses.
- Develop quality data solutions and simplify diverse datasets into data models that encourage self-service.
- Build resilient data pipelines optimized for data quality and poor-quality source data.
- Own data mapping, business logic, transformations, and data quality.
- Debug low-level systems and measure and optimize performance on large production clusters.
- Participate in architecture discussions, influence the product roadmap, and own new projects.
- Maintain existing platforms and evolve them toward newer technology stacks and architectures.
- Collaborate with developers, analysts, operations, and other cross-functional partners.
Requirements
- At least 8 years of professional experience as a data engineer.
- BS in Computer Science required; an MS in Computer Science is preferred.
- Extensive SQL skills.
- Proficiency in at least one scripting language, with Python preferred.
- Extensive experience with Hadoop, HDFS, YARN, MapReduce, Hive, Kafka, Spark, Airflow, and Presto or Trino.
- Deep Apache Spark expertise, including performance tuning, optimization, and scalable batch and streaming pipeline development.
- Proficiency in designing, implementing, and optimizing conceptual, logical, and physical data models.
- Experience with AWS, GCP, or Looker is a plus.
- AI literacy or an AI growth mindset.
Benefits
- Hybrid work with Monday through Thursday in the office and flexible remote work on Fridays.
- Healthcare options including medical, dental, and vision, where locally available.
- Life, accident, disability, commuter, and retirement benefits such as 401(k) or pension, subject to location.
- Global mental health and financial wellness support.
- Paid time off in accordance with local leave policies.
- Reasonable workplace accommodations and adjustments.
Tech Stack
Apache AirflowApache HadoopApache HiveApache KafkaApache SparkAWSGoogle Cloud PlatformPrestoPythonSQLYarn
Categories
Data Engineering
About Roku
Roku builds streaming players, Roku-branded TVs and audio gear, and the Roku OS licensed to TV manufacturers. Its platform supports ad-supported and subscription streaming, including The Roku Channel and Roku Originals, and underpins a significant advertising business. Founded in 2002 and headquartered in San Jose, it is a public company trading on NASDAQ under the ticker ROKU.
