6 months ago
Remote, IndiaSenior
Responsibilities
- Design, code, test, and deliver software that automates manual operational work.
- Troubleshoot priority incidents, facilitate blameless post-mortems, and ensure permanent incident resolution.
- Collaborate with development teams throughout the software life cycle to improve reliability and scale.
- Identify application patterns and analytics that support improved service-level objectives.
- Design self-healing and resiliency patterns.
- Automate software and product upgrades, change management, and release management.
- Participate in 24×7 support coverage as needed.
- Mentor and guide junior developers.
Requirements
- Expertise in at least one technology stack for designing, coding, testing, and delivering software.
- Proficiency in one or more technology domains, including the ability to solve complex and mission-critical problems.
- Working knowledge of infrastructure components such as routers, load balancers, cloud products, container systems, compute, storage, and networks.
- Excellent debugging and troubleshooting skills.
- Prior experience with DevOps and/or application development teams.
- Hands-on large-scale software development experience, preferably with Java, Python, or scripting languages.
- Hands-on experience with Kubernetes, Docker, and Docker Swarm-style deployments.
- Exposure to Datadog monitoring.
- Hands-on experience with continuous delivery tools.
- Hands-on experience with Unix systems, including Linux and Solaris.
- Exposure to orchestration and configuration management tools for applications.
- Experience with infrastructure components used in data warehousing or big data environments.
- Excellent written and oral communication skills for senior technical and business audiences.
- Ability to prioritize effectively in a dynamic, globally focused work environment.
Tech Stack
Categories
Site Reliability
