4 months ago
Bengaluru, IndiaStaff+
Responsibilities
- Define and evolve the end-to-end architecture for core platform services, distributed systems, service orchestration, data flows, and enterprise workflows.
- Drive architectural reviews, RFCs, long-term technical planning, platform contracts, service boundaries, and internal engineering standards.
- Set technical direction for cloud-native compute, storage, networking, deployment models, containerization, orchestration, and infrastructure as code.
- Design and build foundational APIs, frameworks, abstractions, and modular platform components that enable application and AI engineers.
- Own production readiness and operational excellence for mission-critical components, including availability, latency, scalability, fault isolation, observability, alerting, and incident response.
- Define and enforce SLAs, SLOs, reliability standards, and operational maturity practices.
- Mentor Staff and Senior engineers and provide technical leadership across engineering, AI research, product, security, and design teams.
Requirements
- 12+ years of experience designing, building, and operating large-scale, production-grade software systems.
- Proven ownership of complex distributed systems or platform architectures, with deep expertise in cloud-native systems, service-oriented design, and distributed architectures.
- Strong hands-on experience with Docker, Kubernetes, and modern cloud platforms.
- Expert understanding of software design principles, clean code practices, system-level trade-offs, and end-to-end software development and operations.
- Strong proficiency in Python and deep experience with at least one of Java, Go, or JavaScript.
- Experience with Terraform, CloudFormation, or equivalent infrastructure-as-code tools.
- Strong understanding of observability, metrics, logging, and distributed tracing.
- Experience with event-driven systems, streaming platforms such as Kafka, and SQL and NoSQL databases.
- Familiarity with enterprise security, access control, and compliance considerations.
- Experience designing cross-system workflows and hands-on experience with BPMN-based or workflow-driven systems.
- Experience supporting AI/ML or data-intensive workloads, MLOps, AI service deployment, or production integration of LLMs, agents, retrieval, or evaluation pipelines is a strong plus.
Benefits
- Opportunity to build foundational AI-native enterprise systems with significant technical ownership and company-wide influence.
- Headquartered in Los Altos, California.
- The posting does not state additional benefits, work arrangement, or compensation.
