
Sr. Site Reliability Engineer - Core Platform & Embedded Reliability (Hybrid)
CrowdStrike1 day ago
Base Salary
$140k - $215k/yr
Responsibilities
- Partner with engineering leadership across product groups to define and execute multi-year reliability roadmaps.
- Design and implement architectural improvements to services, libraries, and platforms used across CrowdStrike.
- Develop and maintain highly reliable and scalable distributed services and shared libraries.
- Lead reliability, scalability, performance, cost optimization, observability, and resilience engineering initiatives.
- Define service-level objectives and error budgets and use them to guide prioritization and decision-making.
- Build infrastructure-as-code and automation that improve reliability and eliminate manual toil.
- Conduct chaos experiments, failure injection, failure modeling, and graceful-degradation design.
- Provide technical leadership during complex incidents and ensure retrospectives result in durable improvements.
- Extract common patterns into shared libraries and tools and collaborate with platform teams on organization-wide improvements.
- Mentor engineers, establish architectural standards, influence technical decisions, and contribute to open-source software and Go best practices.
Requirements
- 10+ years of experience building and operating distributed systems and service-oriented backends at scale.
- 5+ years developing SaaS microservices using a modern backend language such as Go, Java, Scala, Kotlin, Python, or Node.js.
- Expert-level proficiency in at least one programming language, with expert Go proficiency or the ability and willingness to reach that level.
- Deep knowledge of distributed systems, consensus, replication, consistency models, failure modes, scalability, concurrency, and parallel processing.
- Experience making and delivering architectural decisions at organizational scope and influencing across teams without direct authority.
- Strong command of engineering best practices, testing, peer code review, and resilient architecture.
- Degree in Computer Science or commensurate experience in data structures, algorithms, and distributed systems.
- Experience using AI technologies to improve decision-making, workflows, efficiency, and business outcomes.
- Preferred experience with large microservice organizations, Kubernetes or similar orchestration systems, AWS, Cassandra, Kafka, Elasticsearch/OpenSearch, Google Cloud Platform, Oracle Cloud Infrastructure, multi-cloud systems, internal platforms, cost optimization, performance engineering, chaos engineering, SLO/SLI frameworks, open-source contributions, or cybersecurity and intelligence.
Benefits
- Market-leading compensation and equity awards, comprehensive physical and mental wellness programs, vacation and holidays, paid parental and adoption leave, professional development, employee networks, volunteer opportunities, and office amenities.
- Hybrid work arrangement requiring typically 2 to 3 days per week in the office.
- Role locations are limited to Midtown Manhattan, NY; Redmond, WA; Sunnyvale, CA; or Austin, TX.
Tech Stack
Apache CassandraApache KafkaAWSElasticsearchGoGoogle Cloud PlatformJavaKotlinKubernetesNode.jsOracle CloudPythonScala
Categories
Site Reliability
About CrowdStrike
CrowdStrike builds the Falcon cloud-native security platform used by enterprises and governments to protect endpoints, cloud workloads, identities, and data with EDR, next-generation antivirus, threat intelligence, and managed threat hunting. Founded in 2011 and publicly traded on NASDAQ as CRWD, the company sells primarily by subscription and also provides incident response services; its systems ingest and analyze nearly 3 trillion security events per day across customer environments.