General Motors

Staff Engineer, Site Reliability Engineering

General Motors
Apply
2 days ago
Markham, CanadaStaff+
H1B Sponsor

Base Salary

$147k - $197k/yr

Responsibilities

  • Lead the design and implementation of scalable, fault-tolerant, and observable infrastructure for vehicle telemetry, data ingestion, and platform operations.
  • Lead production-readiness initiatives across multiple teams through hands-on coding, reliability standards, architectural improvements, observability, and resilient deployments.
  • Design and improve CI/CD pipelines with quality gates, artifact promotion, deployment verification, progressive delivery, and rollback practices.
  • Define and implement SLOs, SLIs, monitoring, alerting, runbooks, and operational best practices.
  • Build reusable AI workflows, skills, and evaluations using regression testing, structured evaluations, and LLM-as-a-judge techniques where appropriate.
  • Automate incident intake, triage, diagnostics, remediation, evidence collection, service requests, and customer-facing status workflows.
  • Participate in a weekly on-call rotation with 12-hour shifts every eight weeks and lead production incident response.
  • Drive post-incident reviews and durable system-level reliability improvements.
  • Partner with internal customers, communicate technical trade-offs, mentor engineers, and influence technical direction through reviews, documentation, and cross-functional projects.
  • Balance reliability, performance, security, delivery speed, and cost in technical decisions.

Requirements

  • 8+ years of experience in SRE, DevOps, or systems engineering, including experience managing or mentoring high-impact teams.
  • Track record designing, building, and maintaining high-scale cloud-native production systems, preferably on Azure, AWS, or GCP.
  • Hands-on experience with observability instrumentation, OTEL collector configuration, SLO/SLI definitions, monitors, alerts, and dashboards.
  • Strong understanding of production readiness, service ownership, incident management, post-incident learning, and continuous reliability improvement.
  • Experience participating in on-call rotations and leading technical responses to production incidents.
  • Experience designing, operating, and improving CI/CD pipelines, including GitOps, release strategies, quality gates, deployment verification, progressive delivery, and safe rollback.
  • Strong programming ability in Python, Go, Java, or a comparable language, with disciplined code review, version control, testing, and maintainability practices.
  • Familiarity with AI-assisted software development, LLM application practices, agentic workflows, and evaluation techniques.
  • Ability to work professionally and empathetically with internal customers in difficult or high-pressure situations.
  • Ability to influence without formal authority and drive adoption of shared engineering patterns.
  • Strong written and verbal communication skills for technical and non-technical audiences.
  • BS, MS, or PhD in computer science, engineering, physics, mathematics, or another relevant technical field.
  • Preferred experience with Azure Databricks, Azure Event Hubs, Azure Kubernetes Service, Helm, Kustomize, Terraform, GitHub Actions, Argo CD, Prometheus, Grafana, Datadog, OpenTelemetry, Promptfoo, CoPilot, Fivetran, Apache Flink, Kafka, or Pulsar.
  • Experience with vehicle telemetry, connected-vehicle platforms, or other high-volume event-driven systems is preferred.

Benefits

  • Hybrid work arrangement requiring attendance at the Markham office at least three times per week.
  • Paid time off, vacation days, holidays, and supplemental pregnancy, parental, and adoption leave benefits.
  • Healthcare, dental, vision, and life insurance benefits.
  • Company and matching contributions to a defined contribution pension plan.
  • GM Vehicle Purchase Plan for employees and their families.
  • The role does not provide immigration-related sponsorship.

Tech Stack

Apache FlinkArgo CDAWSAzureDatabricksDatadogGitHub ActionsGoGoogle Cloud PlatformGrafanaHelmJavaKubernetesPrometheusPythonTerraform

Categories

DevOpsSite Reliability
General Motors

About General Motors

10,000+ employees

General Motors’ vision is to create a world with Zero Crashes, Zero Emissions and Zero Congestion, and we have committed ourselves to leading the way toward this future. Today, we are in the midst of a transportation revolution, and we have the ambition, the talent and the technology to realize the safer, better and more sustainable world we want. As an open, inclusive company, we’re also creating an environment where everyone feels welcomed and valued for who they are. One team, where all ideas are considered and heard, where everyone can contribute to their fullest potential, with a culture based in respect, integrity, accountability and equality. Our team brings wide-ranging perspectives and experiences to solving the complex transportation challenges of today and tomorrow. For information on the GM Privacy Statement, please visit http://www.gm.com/privacy-statement.html

Contact me