Engineer III - AI Powered SRE (UK Shift Timing)
CrowdStrikeResponsibilities
- Design and develop AI agents that detect, diagnose, and remediate operational issues across bare-metal servers and virtual machines.
- Document operational workflows and build LLM-powered, event-driven automation pipelines using agentic frameworks or custom solutions.
- Integrate agentic workflows with ELK, Prometheus, Grafana, and other monitoring and telemetry systems for anomaly detection and incident response.
- Develop AI-assisted incident response tooling that correlates distributed-system signals and generates remediation recommendations.
- Support platform availability, latency, throughput, monitoring, and issue response through operational automation.
- Instrument and evaluate agent performance and establish feedback loops to improve accuracy and reliability.
- Participate in incident analysis and post-incident reviews using AI tooling to accelerate resolution.
- Gather and analyze operating-system and application metrics for AI models and predictive operations.
- Write technical documentation for automation systems, agent workflows, and operational runbooks.
- Collaborate with globally distributed engineers and senior technical staff on shared automation solutions.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.
- At least 5 years of software engineering experience with exposure to systems and infrastructure concepts.
- At least 5 years of hands-on experience with one or more of Python, Go, Java, or C++, with Python strongly preferred.
- Foundational experience or strong interest in AI agents or LLM-powered automation, including prompt engineering, tool or function calling, and retrieval-augmented generation.
- Familiarity with agentic frameworks such as LangChain, LlamaIndex, AutoGen, or CrewAI.
- Experience integrating automation with ticketing systems, monitoring platforms, or operational pipelines.
- Working knowledge of Linux engineering and administration.
- Familiarity with Linux, VMware, Docker, Kubernetes, or related infrastructure platforms, including solid understanding of Kubernetes.
- Familiarity with ELK, Prometheus, or Grafana monitoring and telemetry stacks.
- Familiarity with Puppet, Chef, Ansible, or equivalent configuration-management tools.
- Understanding of software engineering fundamentals and willingness to learn distributed-systems concepts.
- Strong analytical, problem-solving, communication, collaboration, ownership, and curiosity skills.
- Proven experience using AI technologies to improve decision-making, workflows, efficiency, and business outcomes.
- Preferred qualifications include MLOps or AI observability experience, SRE, Platform Engineering, or DevOps exposure, self-healing infrastructure experience, and contributions to open-source AI or SRE tooling.
Benefits
- Market-leading compensation and equity awards are offered, though no base salary amount is stated.
- Comprehensive physical and mental wellness programs are provided.
- Competitive vacation and holidays are provided.
- Paid parental and adoption leave is available.
- Professional development opportunities are available to all employees.
- Employee Networks, geographic neighborhood groups, and volunteer opportunities support connection and community.
- The company offers a vibrant office culture with world-class amenities.
- CrowdStrike is Great Place to Work Certified globally.
- The role follows UK shift timing and is part of a globally distributed team.
Categories
About CrowdStrike
CrowdStrike (Nasdaq: CRWD), a global cybersecurity leader, has redefined modern security with the world’s most advanced cloud-native platform for protecting critical areas of enterprise risk — endpoints and cloud workloads, identity and data. Powered by the CrowdStrike Security Cloud and world-class AI, the CrowdStrike Falcon® platform leverages real-time indicators of attack, threat intelligence, evolving adversary tradecraft and enriched telemetry from across the enterprise to deliver hyper-accurate detections, automated protection and remediation, elite threat hunting and prioritized observability of vulnerabilities. Purpose-built in the cloud with a single lightweight-agent architecture, the Falcon platform delivers rapid and scalable deployment, superior protection and performance, reduced complexity and immediate time-to-value. CrowdStrike: We stop breaches.