3 months ago
Responsibilities
- Own end-to-end architecture and system design for large, complex ML Platform projects.
- Define technical direction and long-term strategy for next-generation ML infrastructure.
- Design solutions for low-latency model inference, large-scale feature stores, real-time monitoring, and LLM/agent orchestration.
- Lead high-impact projects from requirements through design, implementation, and production operation.
- Translate the needs of ML engineers, data scientists, and product teams into functional requirements and scalable technical solutions.
- Arbitrate technical decisions while balancing latency, reliability, cost, and security constraints.
- Represent engineering and advise senior leadership on technical considerations across the ML lifecycle.
- Drive cross-team initiatives that improve ML development velocity and MLOps maturity.
- Mentor and grow engineers while modeling strong software design, implementation, and operations practices.
Requirements
- 10+ years of professional software development experience or equivalent domain expertise, with a strong background in service-oriented architecture and large-scale distributed systems.
- Demonstrated experience serving as a technical lead, providing technical direction, leading multi-team initiatives, and mentoring engineers.
- Experience working on production machine learning platform services.
- Strong product instincts and understanding of business context.
- Strong communication skills with technical and non-technical stakeholders.
- Ability to collaborate cross-functionally with ML engineers, data scientists, software engineers, product managers, and business stakeholders.
- Ability to operate autonomously and effectively in ambiguous environments.
- Hands-on experience using AI tools to accelerate work.
- Preferred: experience building large-scale ML serving or data infrastructure, including model inference, feature stores, real-time feature computation, or model registries.
- Preferred: familiarity with LLMs, LLM frameworks, and agentic AI patterns such as tool use, multi-agent orchestration, and retrieval-augmented generation.
- Preferred: familiarity with AWS and cloud-based AI/ML services such as SageMaker, Bedrock, Databricks, and OpenAI.
- Preferred: experience training and shipping machine learning models to production.
- Preferred: ability to synthesize ideas across an organization and establish technical vision.
- Preferred: comfort working with geographically distributed teams.
- Preferred: passion for side projects, open source, or self-directed technical initiatives.
Tech Stack
AWSDatabricks
Categories
About Stripe
Stripe builds programmable financial services. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Headquartered in San Francisco and Dublin, the company aims to increase the GDP of the internet.