
Site Reliability Engineer
OneStream Software3 months ago
Remote, United StatesSenior
Base Salary
$114k - $148k/yr
Responsibilities
- Implement application and infrastructure observability solutions for availability, reliability, and performance.
- Participate in on-call rotations, incident response, post-mortems, and review meetings.
- Partner with Product and Engineering teams to develop, deploy, and maintain reliable systems and services.
- Influence designs, architectures, standards, and methods for large-scale systems.
- Automate processes to improve reliability, performance, and availability.
- Maintain technical documentation, workflows, and knowledge base articles.
- Provide pull request feedback and participate in peer coding reviews.
- Build integrations between Dynatrace, Azure DevOps, and Jira using codified automation.
- Apply OneStream Software knowledge and SOC/FedRAMP controls with Compliance and Security teams.
- Mentor others across several technical areas.
Requirements
- Bachelor’s degree in computer science, engineering, or a technology-related field, or equivalent work experience.
- Proven experience as a Site Reliability Engineer or in a similar role.
- At least 6 years of cloud infrastructure and software development experience.
- At least 2 years of hands-on Azure Kubernetes Service experience with container-based deployments or comparable platforms.
- Advanced knowledge of application performance monitoring and observability tools.
- Advanced knowledge of infrastructure-as-code concepts and tooling across Azure, AWS, or GCP.
- Deep knowledge of configuration management and orchestration tools including Ansible, PowerShell DSC, Chef, and Puppet.
- At least 6 years of hands-on experience automating with PowerShell, Bash, CLI, REST APIs, Python, ARM templates, or other scripting languages.
- Experience with source control tools such as Git, Azure DevOps, or GitHub.
- Knowledge of Kubernetes, OpenShift, AKS, GKS, or Helm.
- Preferred experience with cloud, managed service, or SaaS providers; Azure infrastructure-as-code deployments; Microsoft and .NET technologies; Linux distributions; and reliable containerized applications.
- Ability to work independently, handle ambiguity, prioritize work, communicate across management and engineering levels, and solve emerging problems.
Benefits
- Remote, United States work arrangement with full-time employment.
- Medical, dental, vision, life insurance, short- and long-term disability, vacation time, and paid holidays.
- Retirement plan and professional development opportunities.
- Travel is not expected to exceed 5%.
- Candidates must be legally authorized to work in the United States without sponsorship.
Tech Stack
AnsibleAWSAzureBashC#ChefDatadogGitGoogle Cloud PlatformGrafanaHelmKubernetesLinux.NETOpenShiftPowerShellPrometheusPuppetPythonSQLTerraform
Categories
Site Reliability