
Staff Engineer, Machine Learning Systems & Reliability - Moveworks
ServiceNow19 hours ago
Responsibilities
- Design and build production systems covering the full ML lifecycle from data and feature preparation through training, evaluation, serving, monitoring, feedback collection, and retraining.
- Build continuous-delivery workflows and automated quality, safety, performance, and compatibility checks for models, prompts, agent workflows, data dependencies, and supporting services.
- Implement shadow traffic, canary releases, progressive delivery, feature flags, versioned artifacts, automated rollback, and operational kill switches.
- Create controlled feedback loops with lineage, validation, candidate comparison, model refresh triggers, and guarded promotion of changes.
- Define and operate SLIs, SLOs, alerts, and error budgets across infrastructure, data pipelines, inference services, model quality, and product behavior.
- Improve the scalability, availability, latency, and cost efficiency of distributed training, inference, and data-processing workloads, including GPU capacity planning and resource optimization.
- Participate in architecture reviews, deployments, on-call, incident response, blameless postmortems, and systemic remediation.
- Build self-service platforms and automation that reduce operational toil and accelerate ML engineers' and data scientists' path to production.
- Apply LLMs or agentic automation to evaluation, troubleshooting, and operational workflows when they provide reliable, measurable improvements.
- Establish practical standards for cloud infrastructure, Kubernetes, infrastructure as code, observability, security, and compliance.
- Provide technical leadership, mentorship, and architectural influence across ML, data, product, and platform teams.
Requirements
- Typically 7+ years of experience in software engineering, platform engineering, SRE, production engineering, or ML infrastructure, with Staff-level technical ownership.
- Strong software-engineering skills in Python and at least one production systems language such as Go, Java, C++, or Rust.
- Experience designing, operating, and troubleshooting distributed production systems, including failure analysis, capacity planning, and performance optimization.
- Hands-on experience with cloud infrastructure, containers, Kubernetes, infrastructure as code, CI/CD, and modern observability.
- Practical understanding of training, evaluation, model deployment, serving, monitoring, versioning, and retraining across the ML lifecycle.
- Ability to distinguish service-health problems from data-quality and model-quality problems.
- Familiarity with SLIs, SLOs, error budgets, sustainable on-call, incident management, and blameless postmortems.
- Strong automation and internal-customer mindset, with excellent technical judgment and communication skills.
- Experience coordinating across teams during production incidents and navigating ambiguity.
Benefits
- Regular employee position with a required-in-office work persona in the AMS - North America and Canada region.
- ServiceNow describes flexible, remote, and required-in-office work personas assigned according to role and location.
- Equal opportunity employment and reasonable accommodation support are provided.
Categories
About ServiceNow
ServiceNow builds a cloud platform for enterprise digital workflows, covering IT service management, customer service, HR service delivery, security operations, and operations management, plus tools for custom app development. It sells subscription SaaS to large organizations and public-sector agencies to automate processes and connect data across systems. Founded in 2004 and headquartered in Santa Clara, California, ServiceNow is a public company listed on the NYSE under the ticker NOW.