
Sr. Site Reliability Engineer
Blackpoint Cyber18 hours ago
Base Salary
$150k - $187k/yr
Responsibilities
- Design, develop, and maintain scalable infrastructure using Terraform and Terragrunt for automated cloud provisioning.
- Own and optimize the AWS cloud environment for cost efficiency, security, resilience, and high availability.
- Manage Kubernetes clusters and related tooling including Helm, ArgoCD, Istio, and Kustomize.
- Administer and scale Confluent Cloud and Apache Kafka data streaming infrastructure.
- Deploy and maintain Redis for caching and real-time data processing.
- Build and maintain monitoring, alerting, and incident response frameworks using Prometheus, Grafana, Alert Manager, Grafana Cloud, OpsGenie, and PagerDuty.
- Facilitate controlled feature deployments and progressive rollouts with LaunchDarkly and PostHog.
- Partner with software development teams to integrate new services, applications, and features.
- Diagnose complex production system issues and implement solutions that improve performance and uptime.
- Drive continuous improvement in automation tooling, operational processes, reliability, and maintainability.
Requirements
- At least five years of experience in a Senior Site Reliability Engineer role or equivalent, with substantial cloud infrastructure and automation experience.
- Expertise with Terraform and Terragrunt for enterprise-scale infrastructure-as-code deployments.
- Comprehensive experience designing, implementing, and maintaining secure, scalable AWS architectures.
- Hands-on experience with Confluent Cloud and Apache Kafka for distributed data streaming.
- Experience with Redis and Amazon RDS for caching and relational database management.
- Experience with OpenSearch, Elasticsearch, or ChaosSearch for enterprise search and analytics.
- Proficiency designing and implementing monitoring and alerting infrastructure with Prometheus, Grafana, Alert Manager, Grafana Cloud, OpsGenie, or PagerDuty.
- Practical experience with LaunchDarkly or PostHog for feature flag and release management.
- Extensive production Kubernetes administration experience with Helm and ArgoCD, plus working knowledge of Istio and Kustomize.
- Strong production troubleshooting, problem-solving, communication, and collaboration skills, including experience in Agile environments.
- Nice-to-have experience with multi-cloud platforms such as Google Cloud Platform and Microsoft Azure.
- Nice-to-have experience with serverless computing, Jenkins, GitHub Actions, and CI/CD pipelines.
- Software development proficiency in Node.js, Python, and/or Go is preferred.
- Understanding of cloud-native and containerized security frameworks and compliance standards is preferred.
Benefits
- Eligible US employees receive health, vision, dental, and life insurance plans, a 401k plan, discretionary time off, and other perks.
- International employees receive benefits based on local market standards and applicable country requirements.
- Equity participation is available globally, with details varying by location and employment structure.
Tech Stack
Apache KafkaAWSAzureElasticsearchGitHub ActionsGoGoogle Cloud PlatformGrafanaHelmIstioJenkinsKubernetesNode.jsPrometheusPythonRedisTerraform
Categories
DevOpsSite Reliability
About Blackpoint Cyber
Blackpoint Cyber provides managed detection, response, and remediation services and a security platform used by managed service providers and mid-market organizations. Founded in 2014 by former U.S. defense and intelligence operators, it offers 24/7 SOC-backed threat hunting and incident response delivered via the Blackpoint Suite. The privately held, remote-first company sells primarily through MSP partners and raised a Series C in 2023.