
Senior Site Reliability Engineer
TeamViewer1 month ago
Responsibilities
- Design, build, and maintain highly available, secure, and scalable Azure cloud infrastructure for the global SaaS platform.
- Manage, scale, and optimize Kubernetes-based containerized platforms using GitOps workflows.
- Deploy feature and maintenance releases.
- Develop monitoring, alerting, and automation systems to support 24/7 service reliability.
- Improve infrastructure through automation and Infrastructure-as-Code.
- Participate in on-call rotations, troubleshoot incidents, and resolve operational issues.
- Identify and mitigate reliability risks while driving innovation and continuous improvement.
Requirements
- Degree in Computer Science, Software Engineering, IT, or equivalent practical experience.
- 5+ years of experience in SRE, DevOps, or software development roles.
- Strong experience with Microsoft Azure; AWS or GCP experience is a plus.
- Proficiency with Infrastructure-as-Code and automation tools such as Terraform and Argo CD.
- Familiarity with programming languages and scripting, including PowerShell.
- Hands-on experience with containers and orchestration tools such as Docker and Kubernetes.
- Knowledge of databases such as MS SQL and Postgres.
- Experience with monitoring and observability tools such as Datadog, Grafana, and Prometheus.
- Understanding of distributed systems, parallel programming, test automation, and network security principles.
- Proactive mindset, eagerness to learn, and strong collaboration skills.
Benefits
- Competitive compensation and bonuses.
- Flexible PTO and paid holidays.
- 401(k) with employer matching.
- Comprehensive health insurance, including 100% employer-paid medical coverage.
- Up to 12 weeks of parental leave.
- Basic life insurance and short- and long-term disability coverage, 100% employer-paid.
- Quarterly team-building events, leadership luncheons, and companywide All Hands meetings.
- Open-door policy and business-casual dress code.
- Work location is Austin, TX.
Tech Stack
Argo CDAWSAzureDatadogDockerGoogle Cloud PlatformGrafanaKubernetesPostgreSQLPowerShellPrometheusTerraform
Categories
Site Reliability