Staff Engineer - Big Data
Bloom Energy2 hours ago
Bengaluru, IndiaStaff+
Responsibilities
- Design, develop, and maintain scalable data pipelines using PySpark and Apache Spark.
- Process and analyze large-scale structured and unstructured datasets in distributed environments.
- Build real-time analytics solutions for cloud and edge devices.
- Solve complex data, architecture, performance-tuning, and debugging problems.
- Implement data wrangling, transformation, and processing solutions.
- Collaborate cross-functionally with data scientists, data engineering teams, and firmware controls teams.
Requirements
- At least 3 years of relevant experience.
- Strong Java or Scala programming and debugging ability with understanding of design patterns.
- Python experience is a bonus.
- Understanding of Kafka, Spark, Flink, Hadoop, and HBase internals, with hands-on experience in one or more preferred.
- Demonstrated experience working with large datasets and implementing data processing solutions.
- Experience tuning and debugging Spark jobs.
- Good understanding of distributed computing principles.
- Knowledge of AWS, GCP, or Azure is beneficial.
- Exposure to data lakes, data warehousing concepts, SQL, and NoSQL databases.
- REST APIs and gRPC experience are beneficial.
- MTech or M.S. with emphasis in computational or decision sciences is preferred.
- Ability to adapt quickly to new technologies, concepts, approaches, and environments.
- Strong problem-solving, analytical, learning, and improvement mindset.
Tech Stack
Apache FlinkApache HadoopApache HBaseApache KafkaApache SparkAWSAzureGoogle Cloud PlatformgRPCJavaPythonScalaSQL
Categories
Data Engineering