17 days ago
Responsibilities
- Own end-to-end architecture and system design for large, complex ML Platform projects.
- Define technical direction and long-term platform strategy for ambiguous, high-impact initiatives.
- Design systems for ML workflow orchestration, CPU and GPU compute, model training, LLM fine-tuning, low-latency inference, feature stores, monitoring, and agent orchestration.
- Lead projects from requirements through design, implementation, and production operation.
- Translate the needs of ML engineers, data scientists, and product teams into scalable technical solutions.
- Arbitrate decisions involving latency, reliability, cost, and security constraints.
- Advise senior leaders on technical considerations across the end-to-end ML lifecycle.
- Drive cross-team initiatives to improve ML development velocity and MLOps maturity.
- Mentor and grow engineers while modeling strong software design, implementation, and operations practices.
Requirements
- 10+ years of professional software development experience or equivalent domain expertise.
- Background in service-oriented architecture and large-scale distributed systems.
- Experience serving as a technical lead, providing technical direction, leading multi-team initiatives, and mentoring engineers.
- Experience building and operating production ML platforms for model training, model serving, orchestration, or ML data systems.
- Experience addressing performance, reliability, scalability, and cost-efficiency requirements.
- Strong product instincts and understanding of business context.
- Strong communication and cross-functional collaboration skills with technical and non-technical stakeholders.
- Ability to work autonomously and effectively in ambiguous environments.
- Hands-on experience using AI tools to accelerate work.
- Preferred: experience with large-scale ML training, serving, data infrastructure, distributed training, accelerator-backed compute, feature stores, real-time feature computation, model registries, experiment tracking, and model evaluation.
- Preferred: experience training and shipping machine learning models to production.
- Preferred: familiarity with LLMs, LLM application frameworks, agentic AI patterns, tool use, multi-agent orchestration, and retrieval-augmented generation.
- Preferred: familiarity with AWS and cloud-based AI and ML services such as SageMaker, Bedrock, Databricks, and OpenAI.
- Preferred: experience rapidly prototyping, iterating from user feedback, setting technical vision, working with distributed teams, or contributing to side projects and open source.
Tech Stack
AWSDatabricks
Categories
About Stripe
Stripe builds programmable financial services. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Headquartered in San Francisco and Dublin, the company aims to increase the GDP of the internet.
