Sight Machine

Staff Site Reliability Engineer

Sight Machine
Apply
21 days ago
Remote, United StatesStaff+

Responsibilities

  • Drive reliability, automation, scalability, and technical direction across Sight Machine’s cloud infrastructure.
  • Design, build, and operate infrastructure for agentic AI workloads, LLM gateways, agent orchestration, and non-deterministic services.
  • Evolve SLO, error-budget, incident-response, postmortem, and reliability-review practices.
  • Troubleshoot complex issues spanning CI/CD, container orchestration, networking, operating systems, cloud resources, databases, and AI/LLM orchestration.
  • Build monitoring, alerting, observability, operational runbooks, automation, internal platforms, and developer tooling.
  • Participate in on-call coverage, improve escalation paths, and reduce avoidable pages through automation.
  • Mentor senior and mid-level engineers and influence architecture decisions across teams without formal authority.
  • Lead self-directed cross-team initiatives that improve stability, reliability, and availability.

Requirements

  • 10+ years of experience with Kubernetes and Docker in at least one of Azure, GCP, or AWS, including production-scale multi-tenant or multi-cluster environments.
  • 10+ years of coding experience with Python, Go, Java, or similar, including building tools and platforms used by other engineers.
  • 10+ years of experience with infrastructure as code and CI/CD tooling such as Terraform or OpenTofu, FluxCD or similar GitOps tooling, and Jenkins or GitHub Actions.
  • Demonstrated production experience designing, building, or operating agentic AI or LLM-based systems.
  • Strong Linux and TCP/IP and application-layer networking fundamentals.
  • Practical experience integrating or operating production LLM gateways, API-based orchestration, or agent frameworks.
  • Operational experience with monitoring and alerting systems such as Prometheus, Grafana, Loki, Sentry, and Signoz or equivalents.
  • Deep understanding of cloud performance and the ability to diagnose complex bottlenecks.
  • Experience authoring technical documentation, including design documents, ADRs, and runbooks.
  • Demonstrated mentorship experience without formal management authority and the ability to influence architecture decisions across teams.
  • Strong hands-on judgment under pressure, including balancing technical risk with customer impact.
  • Preferred experience with Kubernetes, FluxCD, Terraform, Helm Charts, Prometheus, Elasticsearch, Python, Java, Kafka, Postgres, and Jenkins.
  • Interest or experience in industrial IoT, analytics, or manufacturing is a plus.

Benefits

  • Competitive salary and stock options.
  • Health care coverage, life insurance, health savings account, and flexible spending account, including spouse and children.
  • Flexible vacation policy.
  • Adaptable working schedule and environment.
  • Hybrid work flexibility, with the ideal candidate near the San Francisco or Ann Arbor office; exceptional fully remote candidates may be considered.
  • Catered lunches, snacks, and beverages.
  • Commuter Savings Program.
  • Company outings.
  • Designated volunteering hours and group volunteer events.

Categories

DevOpsSite Reliability
Sight Machine

About Sight Machine

51-200 employees

Sight Machine provides a manufacturing data platform that unifies plant-floor and enterprise sources to deliver real-time production monitoring, quality analytics, and decision support for industrial manufacturers. It sells subscription software and services to large manufacturers, runs on a standardized infrastructure-as-code deployment model, and is privately held; founded in 2011, it is headquartered in San Francisco with an office in Ann Arbor.

Contact me