
Engineer II - Site Reliability (Hybrid, IND)
CrowdStrikeResponsibilities
- Operate Temporal infrastructure in production across multiple environments using Helm, Kubernetes, and FluxCD.
- Automate deployments, upgrades, scaling, troubleshooting, and other recurring operational work using scripts and workflows.
- Support capacity planning, performance tuning, resource utilization analysis, and bottleneck identification.
- Improve observability through metrics, logs, dashboards, and alerting.
- Participate in incident response and the on-call rotation, including triage, escalation, incident summaries, and runbook creation.
- Manage infrastructure configuration through GitOps workflows and infrastructure-as-code pull requests.
- Troubleshoot deployment, connectivity, and performance issues and help determine preventive fixes.
- Support internal engineering teams with Temporal onboarding, integration debugging, questions, and documentation.
- Work with PostgreSQL, AWS or GCP, Kubernetes networking, Helm charts, certificate rotation, secret management, and distributed systems operations.
Requirements
- At least 3 years of experience in DevOps, SRE, platform engineering, or infrastructure roles.
- Experience deploying and troubleshooting services on Kubernetes, including familiarity with pods, deployments, services, kubectl, and YAML.
- Experience using Helm, including charts, values files, and failed-release troubleshooting.
- Some infrastructure-as-code or declarative infrastructure experience with tools such as Terraform, Ansible, FluxCD, or ArgoCD.
- Exposure to AWS or GCP and basic compute, networking, and storage concepts.
- Ability to write Bash, Python, or Go scripts for automation and simple tooling.
- Foundational experience with databases or other persistent services, preferably PostgreSQL, including backups, schema management, and connection handling.
- Willingness to learn unfamiliar systems and ask for help.
- Proven experience using AI technologies to improve decision-making, workflows, processes, efficiency, or business outcomes.
- Preferred qualifications include experience with Temporal, Airflow, Prefect, Cadence, FluxCD, ArgoCD, distributed tracing, observability platforms, Go, internal platform teams, PostgreSQL performance tuning, and multi-region deployment or failover patterns.
Benefits
- Compensation and equity awards are provided, with the posting describing them as market-leading and competitive.
- Comprehensive physical and mental wellness programs.
- Competitive vacation and holidays.
- Paid parental and adoption leave.
- Professional development opportunities for employees at all levels and in all roles.
- Employee Networks, geographic neighborhood groups, and volunteer opportunities.
- Vibrant office culture with world-class amenities.
- Great Place to Work Certified across the globe.
- Hybrid work arrangement, as indicated in the job title.
Tech Stack
Categories
About CrowdStrike
CrowdStrike (Nasdaq: CRWD), a global cybersecurity leader, has redefined modern security with the world’s most advanced cloud-native platform for protecting critical areas of enterprise risk — endpoints and cloud workloads, identity and data. Powered by the CrowdStrike Security Cloud and world-class AI, the CrowdStrike Falcon® platform leverages real-time indicators of attack, threat intelligence, evolving adversary tradecraft and enriched telemetry from across the enterprise to deliver hyper-accurate detections, automated protection and remediation, elite threat hunting and prioritized observability of vulnerabilities. Purpose-built in the cloud with a single lightweight-agent architecture, the Falcon platform delivers rapid and scalable deployment, superior protection and performance, reduced complexity and immediate time-to-value. CrowdStrike: We stop breaches.