1 month ago
Remote, United StatesStaff+
Responsibilities
- Design and develop multi-threaded asynchronous replication systems with parallel streaming capabilities.
- Build object-level delta replication with checkpointing and resume functionality.
- Develop bucket- and share-level replication controls and replication engines.
- Implement secure data transfer using TLS 1.3 with mutual authentication.
- Build checksum validation and verification pipelines to ensure end-to-end data integrity.
- Design manual failover workflows for disaster recovery scenarios.
- Build and maintain APIs for replication configuration, control, and automation.
- Develop metadata tracking and change detection systems for efficient replication.
- Implement RPO visibility, alerting, and operational insights for replication status.
- Contribute to monitoring dashboards focused on replication health and performance.
- Design systems for high availability, fault tolerance, and scalability.
- Partner with QA teams on performance, resiliency, and scale validation.
- Collaborate with backend, security, and platform teams on end-to-end replication workflows.
- Participate in debugging, production issue resolution, and continuous reliability improvement.
- Provide technical leadership, architectural guidance, and mentorship to the engineering team.
Requirements
- 8+ years of experience in distributed systems, storage systems, or backend software engineering.
- Strong programming skills in one or more of C++, Go, Java, or Rust.
- Experience designing and building data replication systems, data pipelines, or distributed data services.
- Deep understanding of consistency, availability, scalability, and fault tolerance in distributed systems.
- Strong expertise in multi-threading, concurrency, and parallel processing.
- Knowledge of TCP/IP, HTTP/HTTPS, TLS, and secure communication.
- Experience implementing checksums, validation, and consistency checks for data integrity.
- Experience designing and building REST APIs and service-based architectures.
- Familiarity with checkpointing, failure recovery, and retry mechanisms in distributed systems.
- Basic understanding of metrics, logging, and alerting.
- Strong debugging, problem-solving, and system design skills.
- Preferred experience with asynchronous replication, disaster recovery, or backup systems.
- Preferred familiarity with object storage, large-scale data storage, delta encoding, change data capture, or incremental data synchronization.
- Preferred experience building high-throughput, low-latency data movement systems and enterprise-scale data platforms.
- Preferred exposure to mutual TLS, encryption, authentication, performance optimization, and large-scale system tuning.