2 hours ago
Pune, IndiaStaff+
Responsibilities
- Own critical customer case escalations end-to-end through root-cause analysis, mitigation, and resolution.
- Lead technical responses for Infinia incidents, including war rooms, live incident bridges, and cross-functional collaboration.
- Troubleshoot complex system, protocol, and application issues across enterprise and hyperscale environments.
- Develop expertise in Infinia internals, metadata handling, storage fabric interfaces, performance tuning, and AI integration.
- Reproduce customer issues and recommend product improvements, workarounds, and reliability enhancements.
- Create and maintain runbooks, performance-tuning guides, and root-cause-analysis documentation.
- Partner with Field CTOs, Solutions Architects, Sales Engineers, Engineering, QA, and Field teams to support strategic customers.
- Use AI tools, observability, automation, and self-healing mechanisms to accelerate diagnostics and reduce MTTR.
- Translate technical findings into executive-ready summaries and business-impact statements.
- Deliver training to customer support and field engineering teams and participate in post-mortems and executive briefings.
Requirements
- At least 8 years of experience in enterprise storage, distributed systems, or cloud infrastructure support/engineering.
- Deep understanding of file systems and interfaces including S3, POSIX, and NFS, as well as storage performance and Linux kernel internals.
- Ability to script and code using Python, Go, and C++.
- Proven debugging experience at the system, protocol, and application levels using tools such as strace, tcpdump, and perf.
- Hands-on Linux troubleshooting experience.
- Exposure to RDMA, NVMe-oF, or high-performance networking stacks.
- Exceptional communication and executive-reporting skills.
- Experience using AI tools such as log-pattern analysis, LLM-based summarization, and automated RCA tooling to accelerate diagnostics and reduce MTTR.
- Experience with DDN, VAST, Weka, or similar scale-out file systems is preferred.
- Strong scripting or coding ability in Python, Bash, or Go is preferred.
- Familiarity with Prometheus, Grafana, ELK, or OpenTelemetry is preferred.
- Knowledge of replication, consistency models, and data-integrity mechanisms is preferred.
- Exposure to Sovereign AI, LLM model-training environments, or autonomous-system data architectures is preferred.
- Willingness to participate in an on-call rotation for after-hours support is required.
Benefits
- The role includes participation in an on-call rotation to provide after-hours support as needed.