3 months ago
Milpitas, CA, USAStaff+
Responsibilities
- Co-own end-to-end architecture for the Nexus platform across hybrid on-premise and cloud environments.
- Define technical direction and architecture roadmaps balancing velocity, scalability, security, and operational maturity.
- Design production-grade agentic workflows, including tool use, multi-agent coordination, evaluation, and safety patterns.
- Architect the MCP ecosystem, LLM gateway, memory and knowledge layers, observability, and platform applications.
- Establish standards for service design, API contracts, security, identity, observability, and developer experience.
- Partner with InfoSec, Cloud Infrastructure, IAM, Networking, and product engineering teams.
- Lead architecture reviews, enterprise architecture forums, and governance processes including ISAR, STARC, and CAB.
- Coach Staff and Senior engineers through design reviews, code reviews, architecture deep dives, and technical mentorship.
- Identify architectural risks, drive remediation, and improve environment separation, observability, and incident response maturity.
- Evaluate emerging LLM, agentic framework, and AI infrastructure technologies and lead targeted proofs of concept.
Requirements
- Master's or PhD in Artificial Intelligence, Machine Learning, Data Science, Computer Science, or a related field.
- At least 7 years of professional software engineering experience, including at least 5 years in architecture or technical leadership roles.
- Demonstrated impact designing and operating large-scale AI/ML, platform, or distributed systems in production.
- Deep proficiency in Python and strong working knowledge of TypeScript/JavaScript, React, and at least one of Go, Java, or C++.
- Expertise with LangGraph, LangChain, LlamaIndex, PyTorch or TensorFlow, and the Hugging Face ecosystem.
- Strong understanding of LLM internals, transformers, embeddings, RAG architectures, fine-tuning, and evaluation methodologies.
- Production experience with LLM providers and gateways such as Anthropic and OpenAI, plus familiarity with MCP and agentic design patterns.
- Experience designing distributed systems, microservices, REST and GraphQL APIs, event-driven architectures, and high-throughput data pipelines.
- Strong experience with Kubernetes, Docker, and hybrid on-premise and cloud workloads.
- Working knowledge of PostgreSQL, MongoDB, Elasticsearch, Redis or Valkey, and vector databases.
- Background in enterprise security and identity, including OAuth, OIDC, SSO, RBAC, secrets management, and data governance.
- Strong communication, stakeholder influence, product judgment, problem decomposition, debugging, and mentoring abilities.
Benefits
- Paid vacation and sick leave.
- Medical, dental, and vision insurance.
- Life, accident, and disability insurance.
- Flexible spending and health savings accounts.
- Employee assistance program and voluntary benefits including supplemental life and AD&D, legal plan, pet insurance, critical illness, accident, and hospital indemnity coverage.
- Tuition reimbursement, transit benefits, the Applause Program, employee stock purchase plan, and Sandisk Savings 401(k) Plan.
- Eligibility for short-term incentives and, depending on role and performance, long-term incentives or restricted stock units.
- The posting anticipates an application deadline of 08/28/2026, subject to earlier closure.
- Compensation range applicability is limited to roles performed in California, Colorado, New York, or eligible remote locations in those states.
Tech Stack
C++DockerElasticsearchGoGraphQLJavaJavaScriptKubernetesMongoDBPostgreSQLPythonPyTorchReactRedisTensorFlowTypeScript