about 4 hours ago
Responsibilities
- Design, engineer, and build SRE platform systems and capabilities with cutting-edge AI.
- Participate in livesite monitoring rotations and handle escalations.
- Drive availability, scalability, and performance improvements based on livesite learnings.
- Ensure technical deliverables meet or exceed expectations on reliability and performance.
- Onboard other teams onto your platforms by driving outcomes yourself.
- Ship early, seek feedback relentlessly, and iterate fast.
- Drive task planning, estimation, scheduling, and staffing.
- Participate in and influence process improvements across the engineering organization.
Requirements
- Proven track record of architecting and engineering large-scale, distributed applications.
- Experience building complex internal platforms adopted by multiple teams.
- Demonstrated ability to drive adoption of systems through hands-on work.
- Experience with AI-powered applications in production.
- Proficiency in object-oriented languages such as C#, C++, Go, or Python.
- Deep understanding of data structures, algorithms, and cloud programming.
- Experience with service-oriented and microservice-based architectures.
- Familiarity with agile development, CI/CD, and DevOps practices.
- Experience with production Kubernetes infrastructure is a plus.
- Experience with cloud providers and managed services is a must.
- Experience with various database backends.