
Senior Kafka SRE Engineer
Charles Schwab2 days ago
Base Salary
$120k - $155k/yr
Responsibilities
- Build, operate, deploy, and continuously improve highly available Confluent Kafka environments across on-premises and cloud platforms.
- Develop Python automation, operational tooling, observability dashboards, alerting solutions, and automated remediation capabilities.
- Apply AIOps practices including anomaly detection, event correlation, predictive analytics, alert reduction, and AI-driven operational insights.
- Deploy and operate Kafka on Kubernetes using Helm and cloud-native operational practices.
- Automate infrastructure provisioning, configuration management, and platform lifecycle activities with Terraform and Ansible.
- Support Kafka capabilities including Schema Registry, Kafka Connect, ksqlDB, Cluster Linking, MirrorMaker, ZooKeeper, KRaft, and RBAC.
- Perform incident response, root-cause analysis, problem resolution, performance tuning, and post-incident improvement.
- Implement observability solutions, operational metrics, SLOs, dashboards, and actionable monitoring controls.
- Partner with engineers, architects, and infrastructure teams to improve platform health, deployment processes, observability, and reliability practices.
Requirements
- 5–7 years of experience supporting and administering enterprise-scale production Confluent Kafka platforms, including Confluent Platform and Confluent Cloud.
- 5–7 years of experience developing Python automation, operational tooling, observability dashboards, and alerting solutions.
- Hands-on production experience with AIOps, predictive monitoring, intelligent automation, incident management, and automated remediation.
- Experience operating Kafka on Kubernetes, preferably Google Kubernetes Engine, using Helm charts.
- Strong experience with Terraform and Ansible for infrastructure provisioning, configuration management, and platform lifecycle automation.
- Deep expertise in Kafka architecture and the Confluent ecosystem, including producers, consumers, partitions, replication, serialization, consumer groups, performance tuning, and exactly-once processing.
- 3+ years of experience with public cloud technologies, preferably Google Cloud Platform.
- Strong Linux administration, troubleshooting, performance tuning, networking, and distributed-systems experience.
- Experience supporting highly available, fault-tolerant, production-critical platforms and performing incident response and root-cause analysis.
- Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related discipline.
- Preferred qualifications include Confluent certifications such as CCDAK or CCAAK and Google Cloud certifications such as Professional Cloud Architect or Professional Cloud DevOps Engineer.
- Preferred experience includes Confluent for Kubernetes, Confluent Cloud APIs, AIOps platforms, Grafana, InfluxDB, BigQuery, Prometheus, Splunk, Datadog, GitHub Actions, Cloud Build, RabbitMQ, IBM MQ, Solace, and Google Pub/Sub.
- Preferred knowledge includes Site Reliability Engineering practices, SLIs, SLOs, error budgets, reliability engineering, and operational excellence.
Benefits
- The role is intended to be performed on site in the specified location or locations.
- The role is eligible for bonus or incentive opportunities in addition to the salary range.
Tech Stack
AnsibleDatadogGitHub ActionsGoogle BigQueryGoogle Cloud PlatformGrafanaHelmInfluxDBKubernetesLinuxPrometheusPythonRabbitMQSplunkTerraform
Categories
DevOpsSite Reliability
About Charles Schwab
Charles Schwab provides brokerage, banking, and wealth management services for individual investors and independent investment advisors. Its business spans trading platforms, advisory and custody services (Schwab Advisor Services), ETFs and mutual funds, and a U.S. bank offering deposits and lending. Founded in 1971 and headquartered in Westlake, Texas, Schwab is publicly traded on the NYSE (SCHW) and expanded its retail and advisor footprint through the acquisition of TD Ameritrade.