Sight Machine

Staff Site Reliability Engineer

Sight Machine
Apply
5 months ago

Responsibilities

  • Drive reliability, automation, scalability, SLOs, error budgets, postmortems, and reliability reviews across the organization.
  • Design, build, and operate infrastructure for agentic AI workloads, LLM gateways, agent orchestration, and non-deterministic services.
  • Troubleshoot complex cross-layer issues involving CI/CD, containers, networking, operating systems, cloud resources, databases, and AI/LLM orchestration.
  • Build monitoring, alerting, observability infrastructure, internal platforms, developer tooling, runbooks, and operational automation.
  • Participate in on-call coverage, improve escalation processes, reduce avoidable pages, and support incident response.
  • Mentor senior and mid-level engineers and influence architecture decisions across teams without formal authority.
  • Lead self-directed cross-team initiatives that improve stability, reliability, and availability.

Requirements

  • 10+ years of experience with Kubernetes and Docker in at least one of Azure, GCP, or AWS, including production-scale multi-tenant or multi-cluster environments.
  • 10+ years of coding experience with Python, Go, Java, or similar, including tools and platforms used by other engineers.
  • 10+ years of experience with infrastructure-as-code and CI/CD tooling such as Terraform or OpenTofu, FluxCD or similar GitOps tooling, and Jenkins or GitHub Actions.
  • Demonstrated production experience designing, building, or operating agentic AI or LLM-based systems.
  • Strong Linux and networking fundamentals, including TCP/IP and application-layer networking.
  • Operational experience with monitoring and alerting systems such as Prometheus, Grafana, Loki, Sentry, and Signoz or equivalents.
  • Deep cloud-performance troubleshooting experience and the ability to diagnose difficult bottlenecks.
  • Experience authoring technical documentation such as design documents, ADRs, and runbooks.
  • Demonstrated mentorship experience without requiring formal management authority.
  • Clear, empathetic communication skills and the ability to challenge architecture decisions appropriately.
  • Ability to work in the San Francisco office on Mondays and Wednesdays.
  • Experience or interest in industrial IoT, analytics, or manufacturing is preferred.

Benefits

  • Competitive salary and stock options.
  • Health care coverage, life insurance, health savings account, and flexible spending account with spouse and children coverage.
  • Flexible vacation policy and adaptable working schedule and environment.
  • Hybrid work flexibility; the position requires working from the San Francisco office on Mondays and Wednesdays, while exceptional candidates may be considered for 100% remote work.
  • Catered lunches, snacks and beverages, casual dress attire, commuter savings program, company outings, and designated volunteering hours with group volunteer events.

Tech Stack

Categories

DevOpsSite Reliability
Sight Machine

About Sight Machine

51-200 employees

Sight Machine provides a manufacturing data platform that unifies plant-floor and enterprise sources to deliver real-time production monitoring, quality analytics, and decision support for industrial manufacturers. It sells subscription software and services to large manufacturers, runs on a standardized infrastructure-as-code deployment model, and is privately held; founded in 2011, it is headquartered in San Francisco with an office in Ann Arbor.

Contact me