
Sr. Site Reliability Engineer
Blackpoint Cyber2 months ago
Remote, CanadaSenior
Responsibilities
- Design, develop, and maintain scalable infrastructure using Terraform and Terragrunt for automated cloud provisioning and orchestration.
- Own and optimize the AWS cloud environment for cost efficiency, security, resilience, and high availability.
- Manage Kubernetes environments using Helm, ArgoCD, Istio, and Kustomize to support continuous delivery and infrastructure-as-code practices.
- Administer and scale Confluent Cloud and Apache Kafka data streaming infrastructure.
- Deploy, configure, and maintain Redis for caching and real-time data processing.
- Implement monitoring, alerting, and incident response frameworks using Prometheus, Grafana, Alert Manager, OpsGenie, and PagerDuty.
- Facilitate controlled feature deployments and progressive rollouts through LaunchDarkly and PostHog.
- Partner with software development teams to integrate new services, applications, and features into existing infrastructure.
- Diagnose and resolve complex production system issues while maximizing performance and uptime.
- Drive continuous improvement of automation tooling, operational processes, and engineering methodologies.
Requirements
- 5+ years of experience in a Senior Site Reliability Engineer role or equivalent, with substantial cloud infrastructure management and automation experience.
- Expertise with Terraform and Terragrunt for enterprise-scale infrastructure-as-code deployments.
- Comprehensive experience designing, implementing, and maintaining secure, scalable, and resilient AWS architectures.
- Extensive hands-on experience with Confluent Cloud and Apache Kafka for distributed data streaming.
- Experience with Redis and Amazon RDS for caching and relational database management.
- Experience with OpenSearch, Elasticsearch, or ChaosSearch for enterprise search and analytics.
- Experience designing and implementing monitoring and alerting infrastructure with Prometheus, Grafana, Alert Manager, OpsGenie, or PagerDuty.
- Practical experience using LaunchDarkly or PostHog for feature flagging and controlled release management.
- Extensive production Kubernetes administration experience with Helm, ArgoCD, and Istio, plus working knowledge of Kustomize.
- Strong production troubleshooting, problem-solving, communication, and collaboration skills, including experience in Agile environments.
- Preferred experience with Google Cloud Platform, Microsoft Azure, security and compliance frameworks, serverless computing, Jenkins, GitHub Actions, Node.js, Python, or Go.
Benefits
- Eligible US employees receive Health, Vision, Dental, and Life Insurance plans.
- Eligible US employees receive a 401k plan and Discretionary Time Off.
- International employees receive benefits aligned with local market standards and applicable country requirements.
- Equity participation is available globally, with program details varying by location and employment structure.
Tech Stack
Apache KafkaAWSAzureElasticsearchGitHub ActionsGoGoogle Cloud PlatformGrafanaHelmIstioJenkinsKubernetesNode.jsPrometheusPythonRedisTerraform
Categories
DevOpsSite Reliability
About Blackpoint Cyber
Blackpoint Cyber provides managed detection, response, and remediation services and a security platform used by managed service providers and mid-market organizations. Founded in 2014 by former U.S. defense and intelligence operators, it offers 24/7 SOC-backed threat hunting and incident response delivered via the Blackpoint Suite. The privately held, remote-first company sells primarily through MSP partners and raised a Series C in 2023.