2 months ago
Pune, IndiaStaff+
Responsibilities
- Own critical customer case escalations, deep root-cause analysis, and mitigation strategies.
- Serve as a technical escalation point for Infinia incidents, especially production-impacting scenarios.
- Lead war rooms, live incident bridges, and cross-functional incident response.
- Use AI-powered debugging, log analysis, automation, and pattern-recognition tools to accelerate resolution.
- Develop subject-matter expertise in Infinia internals, metadata handling, storage fabric interfaces, performance tuning, and AI integration.
- Reproduce complex customer issues and recommend product improvements or workarounds.
- Create and maintain runbooks, performance-tuning guides, and root-cause-analysis documentation.
- Feed support insights into product development to improve reliability and diagnostics.
- Partner with Field CTOs, Solutions Architects, and Sales Engineers on strategic customer success.
- Translate technical issues into executive-ready summaries and business-impact statements.
- Participate in post-mortems and executive briefings.
- Drive adoption of observability, automation, and self-healing support mechanisms using AI/ML tools.
- Deliver training to customer support and field engineering teams.
Requirements
- 8+ years of experience in enterprise storage, distributed systems, or cloud infrastructure support/engineering.
- Deep understanding of file systems and interfaces including S3, POSIX, and NFS, plus storage performance and Linux kernel internals.
- Scripting and coding experience with Python, Go, and C++.
- Proven debugging ability at system, protocol, and application levels using tools such as strace, tcpdump, and perf.
- Hands-on Linux troubleshooting experience.
- Exposure to RDMA, NVMe-oF, or high-performance networking stacks.
- Exceptional communication and executive reporting skills.