20 hours ago
Pune, IndiaStaff+
Responsibilities
- Define and govern AI platform infrastructure architecture, reference architectures, standards, guardrails, patterns, and roadmaps across cloud and on-premises environments.
- Provide technical direction and design authority to infrastructure engineering teams through solution reviews, design approvals, and delivery oversight.
- Evaluate and guide adoption of GPU platforms, Kubernetes distributions, storage and networking solutions, and observability stacks.
- Partner on on-premises platform development, data center integration, and operational readiness.
- Design and guide GPU infrastructure for AI/ML workloads, including capacity planning, scheduling, tenancy and isolation, performance, reliability, and cost controls.
- Embed identity, encryption, logging, vulnerability management, auditability, security, compliance, and operational controls by design.
- Ensure availability, disaster recovery, monitoring and alerting, SLOs, incident readiness, infrastructure automation, and CI/CD enablement.
- Translate platform requirements into secure, scalable, reusable, maintainable, and supportable infrastructure solutions.
- Build trusted relationships with engineering, product, control, Cyber Security, Data, and Model Risk Management stakeholders.
Requirements
- Proven infrastructure architecture leadership for large-scale platforms, including reference architectures, standards, guardrails, and roadmaps across cloud and on-premises environments.
- Hands-on experience building and delivering platform capabilities across cloud and on-premises environments, with strong production engineering and operational knowledge.
- Knowledge of one or more major cloud platforms, including core services, networking, identity, and security principles.
- Experience with platform engineering practices including APIs, containerisation, Kubernetes, CI/CD, automation, and observability.
- Familiarity with infrastructure as code, policy as code, and reusable platform patterns.
- Experience evaluating and adopting vendor products and services in line with enterprise support, lifecycle, and operational models.
- Strong stakeholder management skills and the ability to translate complex requirements into pragmatic infrastructure designs.
- Understanding of agent, skill, and MCP processing patterns, including deployment architecture, runtime operations, performance, and reliability considerations.
- Familiarity with FinOps and cost optimization for platform infrastructure, particularly GPU and high-throughput data workloads.
- Knowledge of service management, problem management, root-cause analysis, continuous service improvement, and post-incident actions aligned to NFRs and SLOs.
- Knowledge of the GPU ecosystem, including scheduling strategies, partitioning, performance tuning, capacity management, and cost governance.
- Experience defining platform architecture principles, standards, and target-state roadmaps.
Benefits
- Continuous professional development and opportunities to grow within HSBC.
- Flexible working.
- Inclusive and diverse workplace committed to equal opportunity and respect.
- Applications are considered based on merit and suitability to the role.
Tech Stack
Categories
About HSBC
HSBC is a global bank that provides retail, commercial and investment banking, payments, and wealth management to individuals, SMEs, and multinationals. It earns fees, interest income, and trading revenues through branches and digital platforms across major markets. Headquartered in London, it serves over 40 million customers in 58 countries and is listed in London, Hong Kong, New York, and Bermuda.
