10 hours ago
Bengaluru, IndiaStaff+
Responsibilities
- Define architecture and long-term technical direction for large-scale data platforms and distributed processing systems.
- Design scalable backend services and data infrastructure supporting business-critical workloads.
- Build backend services and platform components primarily in Java, with Python for data processing, orchestration, automation, and tooling.
- Design batch and real-time processing architectures using Apache Spark and Apache Flink.
- Establish patterns for data ingestion, transformation, computation, orchestration, storage, and serving.
- Design workflow and dependency-management capabilities using Apache Airflow or related technologies.
- Define operational patterns for large-scale data and backend workloads on Kubernetes.
- Solve distributed-systems challenges involving scalability, state management, fault tolerance, consistency, partitioning, backpressure, resource management, and recovery.
- Improve platform reliability, performance, observability, developer productivity, and infrastructure efficiency.
- Lead architecture reviews, migrations, modernization initiatives, and technical strategy across multiple teams.
- Establish technical standards and reusable platform capabilities for engineering and data teams.
- Mentor senior engineers and influence engineering practices through design reviews, code reviews, and technical guidance.
- Evaluate emerging technologies and make strategic build-versus-buy and architectural decisions.
Requirements
- 10+ years of software engineering experience, including significant experience designing and operating large-scale distributed systems or data platforms.
- Deep expertise in Java and strong software engineering fundamentals.
- Proficiency with Python for data engineering, automation, or platform development.
- Extensive experience designing and building production backend services and distributed systems.
- Deep hands-on experience with Apache Spark and large-scale distributed data processing.
- Strong experience with the Hadoop ecosystem, including HDFS, Hive, and YARN.
- Experience designing and operating real-time or stateful streaming systems using Apache Flink.
- Experience with large-scale workflow orchestration using Apache Airflow or comparable technologies.
- Strong production experience with Kubernetes, containers, and cloud-native application architectures.
- Deep understanding of distributed-systems concepts, including partitioning, replication, consistency, fault tolerance, distributed state, scheduling, resource management, and failure recovery.
- Understanding of batch and streaming architectures and their tradeoffs.
- Experience driving architecture and technical decisions across multiple teams or major platform initiatives.
- Ability to turn ambiguous business or platform requirements into executable technical strategies.
- Track record of mentoring senior engineers and influencing engineering practices beyond an immediate team.
- Preferred: experience with petabyte-scale datasets or billions of events per day.
- Preferred: experience with Apache Iceberg, multi-tenant data platforms, data governance, lineage, data quality, metadata management, lifecycle management, and platform observability.
Tech Stack
Apache AirflowApache FlinkApache HadoopApache HiveApache KafkaApache SparkDockerJavaKubernetesPythonYarn
Categories
BackendData Engineering
About TCGplayer
TCGplayer operates an online marketplace and software tools for buying and selling trading card games and related collectibles, serving hobby shops and individual sellers. It generates revenue through marketplace fees, seller subscriptions, and fulfillment programs such as TCGplayer Direct. An eBay subsidiary, the company provides authentication and logistics services that help stores list inventory at scale and reach buyers across the U.S. and internationally.
