
Senior Site Reliability Engineer ( SRE)
Charles Schwab5 days ago
Base Salary
$130k - $175k/yr
Responsibilities
- Design and implement automation solutions that reduce manual effort and improve operational efficiency and system reliability.
- Improve availability, scalability, observability, deployment processes, rollout validation, capacity planning, and operational excellence for distributed enterprise platforms.
- Develop and maintain operational tooling, monitoring dashboards, alerting frameworks, and proactive monitoring strategies.
- Apply AI/ML-enabled operational capabilities, anomaly detection, predictive alerting, AIOps, and ML-assisted observability.
- Participate in incident response and on-call support while driving long-term reliability improvements.
- Collaborate with engineering, product, scrum, and operations teams to address complex operational challenges and influence technical strategy.
- Support continuous delivery, GitOps, infrastructure automation, and deployment optimization initiatives.
Requirements
- 8+ years of experience supporting and administering enterprise-scale applications, platforms, or infrastructure environments.
- 6+ years of experience developing automation solutions, operational tooling, monitoring dashboards, and alerting frameworks.
- 6+ years of experience applying SDLC practices and driving process improvement initiatives.
- Experience with SRE, production operations, system monitoring, deployment management, operational excellence, and large-scale distributed systems.
- Experience administering Linux and Windows Server environments, including troubleshooting, performance tuning, and system optimization.
- Experience deploying, supporting, configuring, or migrating cloud-based applications and platforms.
- Knowledge of networking fundamentals including DNS, DHCP, firewalls, and routing.
- Proficiency in one or more of Python, Java, PowerShell, Bash, or .NET.
- Experience with relational or NoSQL databases including SQL Server, Oracle, or MongoDB.
- Knowledge of messaging and event-streaming technologies including Kafka, RabbitMQ, IBM MQ, or Solace.
- Experience with observability and monitoring platforms such as Splunk or AppDynamics.
- Experience applying AI/ML-powered operational practices, including anomaly detection, predictive alerting, AIOps, or ML-assisted observability.
- Bachelor’s degree in computer science, information technology, engineering, or a related field.
- Preferred experience includes financial services, Agile environments, AIOps, intelligent automation, ML-driven observability, CI/CD, GitOps, infrastructure automation, containers, Kubernetes, OpenShift, GCP, AWS, or Microsoft Azure.
- Demonstrated ability to influence technical strategy, drive operational improvements, and lead reliability-focused initiatives across teams.
Benefits
- The role is expected to be performed on site in the specified location(s).
- The role is eligible for bonus or incentive opportunities in addition to the salary range.
Tech Stack
Apache KafkaAWSAzureBashGitHub ActionsGoogle Cloud PlatformHarnessJavaJenkinsKubernetesLinuxMicrosoft SQL ServerMongoDB.NETOpenShiftPowerShellPythonRabbitMQSplunk
Categories
Site Reliability
About Charles Schwab
Charles Schwab provides brokerage, banking, and wealth management services for individual investors and independent investment advisors. Its business spans trading platforms, advisory and custody services (Schwab Advisor Services), ETFs and mutual funds, and a U.S. bank offering deposits and lending. Founded in 1971 and headquartered in Westlake, Texas, Schwab is publicly traded on the NYSE (SCHW) and expanded its retail and advisor footprint through the acquisition of TD Ameritrade.