1 day ago
Toronto, CanadaStaff+
Responsibilities
- Set technical direction for model serving, deployment, monitoring, architecture, engineering patterns, and production tradeoffs.
- Build and operate high-throughput, low-latency, observable production ML services.
- Establish engineering standards for testing, CI/CD, observability, alerting, rollback, and on-call practices.
- Own MLOps capabilities including model and feature monitoring, drift and data-quality detection, retraining, promotion workflows, versioning, safe rollout, and incident response.
- Lead multi-quarter projects and influence engineering decisions across teams and stakeholders.
- Scope and prioritize ambiguous cross-team problems with Product Managers and other stakeholders.
- Develop capabilities for AI agents, recommendations, predictive models, multimodal restaurant-content understanding, and partner business insights.
Requirements
- 7+ years of professional software engineering experience, including substantial experience building and operating machine learning systems in production.
- Experience working across more than one organization and applying industry-standard ML system architectures and operating practices.
- Hands-on experience with AWS, GCP, or Azure as a primary model-serving environment, including managed deployment, scaling, and observability services.
- Strong fundamentals in distributed systems, service and API design, concurrency, latency and throughput tradeoffs, testing, and production ownership with on-call responsibility.
- Strong command of Python and proficiency in at least one strongly typed language, preferably Java.
- Demonstrated experience training, serving, and deploying ML models at production scale.
- Experience with production MLOps, including monitoring, data-quality detection, retraining, promotion, versioning, rollout, rollback, and incident response.
- Track record of technical leadership across multi-quarter projects and cross-functional stakeholders.
- Preferred experience serving LLMs in production, including inference infrastructure, GPU utilization, batching, caching, quantization, and latency-cost optimization.
- Preferred applied ML experience in ranking, recommendations, classification, NLP, RAG, or agentic systems.
- Preferred Kubernetes experience at meaningful production scale.
- Preferred experience developing ETL jobs, especially with Spark, or data warehouse infrastructure.
- Preferred familiarity with A/B testing design and analysis and with introducing tooling or platform capabilities and driving adoption.
Benefits
- Hybrid workplace with an expectation of working in the office two days per week.
- Work from almost anywhere for up to 20 days per year.
- Company-paid therapy sessions through SpringHealth and a company-paid Headspace subscription.
- Annual company-wide week off.
- Paid parental leave, generous paid vacation, birthday time off, 20 days of paid time off, and paid volunteer time.
- Development Dollars, leadership development, and access to thousands of on-demand e-learning courses.
- Travel discounts, Employee Resource Groups, private health and dental insurance, life and disability insurance.
- Health benefits, flexible spending account, retirement benefits, paid sick, medical, bereavement, floating and company holidays, and potential annual bonus and equity eligibility.
Tech Stack
Apache AirflowApache KafkaApache SparkAWSAzureDockerElasticsearchFastAPIFlaskGoogle Cloud PlatformGrafanaGraphiteHelmJavaKubernetesMavenMongoDBPostgreSQLPrometheusPythonPyTorchRedisSnowflakeTeamCityXGBoost
Categories
About OpenTable
OpenTable builds an online restaurant reservations marketplace for diners and reservation, guest management, and operations software for restaurants, sold via subscriptions and per-cover fees. Its consumer apps and integrations power bookings across 70,000+ restaurants worldwide and connect with POS and other hospitality systems. Founded in 1998 and headquartered in San Francisco, OpenTable is part of Booking Holdings.
