
Software Engineer II - Streaming
Rivian and Volkswagen Group Technologies25 days ago
Palo Alto, CA, USAMid Level
Base Salary
$105k - $131k/yr
Responsibilities
- Build and maintain scalable real-time streaming services for vehicle, cloud, and operational data.
- Design, develop, and operate Apache Flink applications and event-driven services using Apache Kafka or Redpanda.
- Create scalable streaming architectures covering topic design, partitioning, checkpointing, savepoints, dead letter queues, replay, and failure recovery.
- Develop production-grade applications in Java, Scala, Python, or Go and optimize latency, throughput, resource utilization, and operational cost.
- Build reusable platform components, SDKs, libraries, and self-service capabilities that improve developer productivity.
- Support deployment and operations with Kubernetes, Docker, CI/CD pipelines, and Infrastructure as Code practices.
- Implement observability through metrics, logging, distributed tracing, dashboards, and alerting.
- Monitor and troubleshoot production systems using consumer lag, checkpoint health, logs, metrics, and operational dashboards.
- Partner with platform, data engineering, ML, analytics, and product teams to deliver real-time data solutions.
- Build streaming pipelines supporting AI/ML platforms, online feature engineering, retrieval-augmented generation, and intelligent applications.
- Participate in production support, incident response, root cause analysis, postmortems, and reliability improvements.
- Use AI-assisted development tools for software design, implementation, testing, debugging, documentation, and productivity.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Software Engineering, or a related technical field, or equivalent practical experience.
- At least 3 years of professional software engineering experience building distributed systems and real-time streaming platforms.
- Hands-on Apache Kafka or Redpanda experience covering topics, partitions, consumer groups, replication, delivery semantics, tuning, optimization, and production operations.
- Strong Apache Flink experience with stateful processing, event time, watermarks, windowing, checkpointing, savepoints, parallelism, backpressure, and fault recovery.
- Strong understanding of event-driven architectures, real-time streaming systems, scalability, fault tolerance, consistency models, replication, and performance optimization.
- Proficiency in one or more of Java, Scala, Python, or Go.
- Strong software engineering fundamentals including data structures, algorithms, object-oriented design, and system design.
- Experience building cloud-native applications with Docker and Kubernetes and using AWS, Azure, or GCP.
- Experience debugging production issues across distributed applications and streaming infrastructure.
- Familiarity with REST APIs, gRPC, or event-driven service integration.
- Experience operating large-scale Kafka and Flink deployments, Kafka Connect, Debezium, or Change Data Capture architectures is preferred.
- Preferred experience includes Avro, Protobuf, JSON Schema, Schema Registry, Apache Iceberg, Delta Lake, ClickHouse, Apache Pinot, Spark Structured Streaming, OpenTelemetry, Prometheus, Grafana, Terraform, and Helm.
- Preferred experience includes real-time feature engineering, AI applications, edge-to-cloud or vehicle-to-cloud streaming, incident response, postmortems, and operational excellence.
Benefits
- Full-time employees may be eligible for an annual company performance bonus and equity in the form of Restricted Stock Units, subject to applicable terms.
- Benefits include health coverage, retirement savings, time off, and family planning programs.
- The role is based in Palo Alto, California; the posting does not specify a remote or hybrid work arrangement.
Tech Stack
Apache FlinkApache KafkaAWSAzureClickHouseDockerGoGoogle Cloud PlatformGrafanagRPCHelmJavaKubernetesPrometheusPythonScalaTerraform
Categories
Data Engineering