
Staff Software Engineer, Caching
Judi Health13 days ago
Charlotte, NC, USA +2 moreStaff+
Base Salary
$167k - $251k/yr
Responsibilities
- Design, build, and operate caching infrastructure serving as the data access layer for complex healthcare domain models.
- Own cache design across eviction policy, service topology, invalidation architecture, and operational recovery.
- Define performance and correctness contracts between the caching layer and dependent services.
- Drive organization-wide caching best practices, eliminate anti-patterns, and build shared libraries.
- Diagnose and resolve caching correctness, performance, and stability issues, including leading incident response and post-incident reviews.
- Set the technical bar through design reviews, code reviews, and mentorship of senior engineers.
- Participate in a 24/7 on-call rotation for continuous operational support and rapid incident response.
Requirements
- 10+ years of production engineering experience with deep specialization in distributed systems and caching infrastructure.
- Production experience operating distributed caching systems such as Redis/Valkey or Memcached at cloud scale, including cluster topology, eviction, replication, and failure recovery.
- Experience with multi-tenancy and shared caching infrastructure across independent workloads.
- Deep knowledge of cache consistency, invalidation strategies, and large-scale failure modes including thundering herd, hot key saturation, cache leasing, and dual-write consistency problems.
- Fluency with cache replacement algorithms beyond LRU, including LFU, ARC, and probabilistic eviction, plus multi-tier caching tradeoffs.
- Experience operating cloud-native systems with strong redundancy, resource-efficiency, and resiliency requirements.
- Track record of leading cross-team technical initiatives through influence and durable documentation.
- Strong written and verbal communication and the ability to produce architectural guidance.
- Ability to operate with ambiguity and drive solutions from first principles through production.
- Strong experience supporting 24/7 production on-call rotations and resolving incidents outside standard business hours.
- Preferred: experience in healthcare, pharmacy benefits, or other regulated industries.
- Preferred: familiarity with probabilistic data structures and real-time transactional processing domains such as payments, financial services, insurance, or healthcare.
- Preferred: familiarity with database replication and change data capture for cache invalidation and source-of-truth consistency.
- Preferred: systems-level programming experience in Rust, Go, or C/C++ for performance-critical infrastructure.
- Preferred: proficiency in Python for infrastructure tooling, automation, or service development.
- Preferred: experience with AWS or GCP for managed caching, database, and streaming services.
- Preferred: open-source contributions, technical publications, or conference presentations in distributed systems or infrastructure.
Benefits
- Hybrid work arrangement with 3 days in the office; offices are located in New York City, Denver, Colorado, and the Charlotte, North Carolina area.
- Participation in a 24/7 on-call rotation is required.