2 months ago
Base Salary
$215k - $260k/yr
Responsibilities
- Own the architecture and evolution of large-scale telemetry pipelines covering ingestion, stream processing, time-series and log storage, and query paths
- Design scalable, durable, fair multi-tenant services that address hot shards, high-cardinality data, noisy neighbors, backpressure, retention, compaction, and graceful degradation
- Improve pipeline reliability and data freshness while reducing on-call burden through system design
- Set technical direction for distributed systems work, lead design reviews, and guide major architectural decisions
- Collaborate with product, compute, networking, and platform teams on observability strategy
- Represent the team’s technical position in cross-organizational discussions
- Participate in customer-facing on-call and use incident response to improve system design
- Coach senior and mid-level engineers through design work, code review, and incident response
- Build reusable patterns and frameworks that improve team effectiveness
Requirements
- 8+ years of software development experience with sustained ownership of production distributed systems
- Deep hands-on experience designing and operating distributed systems at scale, including sharding, replication, consistency, load balancing, and concurrency
- Experience with large-scale observability data infrastructure such as time-series databases, log aggregation, streaming pipelines, or distributed tracing backends
- Familiarity with Prometheus, VictoriaMetrics, Loki, OpenTelemetry, Kafka, Vector, or similar technologies
- Strong programming fundamentals in Go or another modern compiled language, with Go strongly preferred
- Comfort working with Kubernetes, microservices, and CI/CD environments
- Experience supporting customer-facing services through on-call rotations
- Ability to drive technical outcomes across team boundaries and communicate with product, support, and engineering teams
- Ability to scope ambiguous problems, identify non-functional requirements, and evaluate customer and business tradeoffs
- Demonstrated ability to mentor engineers through design guidance, code review, and constructive feedback
Benefits
- Competitive compensation and equity packages, including Restricted Stock Units
- Paid time off, paid holidays, and leave of absence programs
- Comprehensive health, dental, and vision insurance
- Employer HSA contributions
- Paid parental leave
- Paid life insurance and short- and long-term disability coverage
- Professional development and tuition reimbursement
- Mental health and wellness support
- Commuter benefits for parking and transit
- Cell phone stipend
- 401(k) plan with company match up to 4% of salary
- Volunteer time off
- Global travel insurance and emergency assistance
- Daily meals allowance
- Additional location-specific perks and programs
Tech Stack
About Crusoe
As the AI factory company, Crusoe is on a mission to accelerate the abundance of energy and intelligence. The company provides a reliable, scalable, cost-effective, and energy-first solution for AI infrastructure. By harnessing large-scale energy resources, building AI-optimized data centers, and delivering an AI cloud platform, Crusoe empowers its customers to build the future faster.
