GrepJob
Judi Health

Staff Software Engineer, Caching

Judi Health
Apply
13 days ago
Charlotte, NC, USA +2 moreStaff+

Base Salary

$167k - $251k/yr

Responsibilities

  • Design, build, and operate caching infrastructure serving as the data access layer for complex healthcare domain models.
  • Own cache design across eviction policy, service topology, invalidation architecture, and operational recovery.
  • Define performance and correctness contracts between the caching layer and dependent services.
  • Drive organization-wide caching best practices, eliminate anti-patterns, and build shared libraries.
  • Diagnose and resolve caching correctness, performance, and stability issues, including leading incident response and post-incident reviews.
  • Set the technical bar through design reviews, code reviews, and mentorship of senior engineers.
  • Participate in a 24/7 on-call rotation for continuous operational support and rapid incident response.

Requirements

  • 10+ years of production engineering experience with deep specialization in distributed systems and caching infrastructure.
  • Production experience operating distributed caching systems such as Redis/Valkey or Memcached at cloud scale, including cluster topology, eviction, replication, and failure recovery.
  • Experience with multi-tenancy and shared caching infrastructure across independent workloads.
  • Deep knowledge of cache consistency, invalidation strategies, and large-scale failure modes including thundering herd, hot key saturation, cache leasing, and dual-write consistency problems.
  • Fluency with cache replacement algorithms beyond LRU, including LFU, ARC, and probabilistic eviction, plus multi-tier caching tradeoffs.
  • Experience operating cloud-native systems with strong redundancy, resource-efficiency, and resiliency requirements.
  • Track record of leading cross-team technical initiatives through influence and durable documentation.
  • Strong written and verbal communication and the ability to produce architectural guidance.
  • Ability to operate with ambiguity and drive solutions from first principles through production.
  • Strong experience supporting 24/7 production on-call rotations and resolving incidents outside standard business hours.
  • Preferred: experience in healthcare, pharmacy benefits, or other regulated industries.
  • Preferred: familiarity with probabilistic data structures and real-time transactional processing domains such as payments, financial services, insurance, or healthcare.
  • Preferred: familiarity with database replication and change data capture for cache invalidation and source-of-truth consistency.
  • Preferred: systems-level programming experience in Rust, Go, or C/C++ for performance-critical infrastructure.
  • Preferred: proficiency in Python for infrastructure tooling, automation, or service development.
  • Preferred: experience with AWS or GCP for managed caching, database, and streaming services.
  • Preferred: open-source contributions, technical publications, or conference presentations in distributed systems or infrastructure.

Benefits

  • Hybrid work arrangement with 3 days in the office; offices are located in New York City, Denver, Colorado, and the Charlotte, North Carolina area.
  • Participation in a 24/7 on-call rotation is required.