
Sr. Site Reliability Engineer
FreedomPay2 months ago
Remote, United StatesSenior
Responsibilities
- Build and maintain understanding of the platform and custom application stack.
- Implement and improve observability strategies, metrics, and AI/ML-driven monitoring across the development lifecycle.
- Design and maintain automated remediation and self-healing workflows for common failure modes.
- Use AI-assisted tools for automation development, anomaly detection, alert correlation, root-cause analysis, postmortems, and runbooks.
- Handle production escalations, participate in an engineering on-call rotation, and support after-hours incidents on a rotational basis.
- Troubleshoot failed scheduled jobs and data-related concerns.
- Improve incident response procedures, documentation, and integrations between monitoring, ticketing, and remediation systems.
- Champion secure and responsible adoption of AI tooling while protecting sensitive data and meeting security, privacy, and PCI obligations.
Requirements
- Bachelor’s degree in Computer Science or equivalent years of relevant experience.
- At least 5 years of hands-on technical experience in highly available, high-throughput, web-based technology environments.
- Expertise with an enterprise APM platform and its AI/ML-driven AIOps capabilities; Dynatrace is strongly preferred, with Datadog or New Relic as comparable alternatives.
- Experience using Anthropic Claude, OpenAI Codex, Azure AI services, Foundry, or Azure SRE Agent for operational and engineering work.
- Proficiency with PowerShell and/or Python for scripting, automation, tooling, and remediation workflows.
- Strong SQL/T-SQL skills and understanding of DNS, HTTP/HTTPS, load balancing, and TCP/IP routing and switching.
- Working knowledge of container orchestration, IaaS/PaaS cloud services, Azure, VMware, and application development processes.
- Preferred experience includes SLI/SLO implementation, enterprise incident management, AIOps or ML-driven production automation, Azure Kubernetes Service, Windows Server/IIS, PagerDuty Process Automation, real-time transaction processing, PCI practices, MLOps, documentation automation, service catalogs, and QA test automation in CI/CD pipelines.
- Strong problem-solving, communication, organizational, ownership, service, and self-directed learning abilities.
Benefits
- Medical, prescription, dental, and vision coverage.
- Life insurance and retirement plans with company match.
- Commission sharing plan and parental and other leave programs.
- Flexible hybrid working environment based in the Philadelphia area; remote arrangements may be considered for exceptional candidates, with occasional Philadelphia travel required.
- Opportunity for upward mobility, professional development, and career growth.
- Full-time salaried position with rotational on-call and after-hours production support.
- Successful completion of a background check and credit check is required.