
Staff Site Reliability Engineer
Sight Machine5 months ago
Responsibilities
- Drive reliability, automation, scalability, SLOs, error budgets, postmortems, and reliability reviews across the organization.
- Design, build, and operate infrastructure for agentic AI workloads, LLM gateways, agent orchestration, and non-deterministic services.
- Troubleshoot complex cross-layer issues involving CI/CD, containers, networking, operating systems, cloud resources, databases, and AI/LLM orchestration.
- Build monitoring, alerting, observability infrastructure, internal platforms, developer tooling, runbooks, and operational automation.
- Participate in on-call coverage, improve escalation processes, reduce avoidable pages, and support incident response.
- Mentor senior and mid-level engineers and influence architecture decisions across teams without formal authority.
- Lead self-directed cross-team initiatives that improve stability, reliability, and availability.
Requirements
- 10+ years of experience with Kubernetes and Docker in at least one of Azure, GCP, or AWS, including production-scale multi-tenant or multi-cluster environments.
- 10+ years of coding experience with Python, Go, Java, or similar, including tools and platforms used by other engineers.
- 10+ years of experience with infrastructure-as-code and CI/CD tooling such as Terraform or OpenTofu, FluxCD or similar GitOps tooling, and Jenkins or GitHub Actions.
- Demonstrated production experience designing, building, or operating agentic AI or LLM-based systems.
- Strong Linux and networking fundamentals, including TCP/IP and application-layer networking.
- Operational experience with monitoring and alerting systems such as Prometheus, Grafana, Loki, Sentry, and Signoz or equivalents.
- Deep cloud-performance troubleshooting experience and the ability to diagnose difficult bottlenecks.
- Experience authoring technical documentation such as design documents, ADRs, and runbooks.
- Demonstrated mentorship experience without requiring formal management authority.
- Clear, empathetic communication skills and the ability to challenge architecture decisions appropriately.
- Ability to work in the San Francisco office on Mondays and Wednesdays.
- Experience or interest in industrial IoT, analytics, or manufacturing is preferred.
Benefits
- Competitive salary and stock options.
- Health care coverage, life insurance, health savings account, and flexible spending account with spouse and children coverage.
- Flexible vacation policy and adaptable working schedule and environment.
- Hybrid work flexibility; the position requires working from the San Francisco office on Mondays and Wednesdays, while exceptional candidates may be considered for 100% remote work.
- Catered lunches, snacks and beverages, casual dress attire, commuter savings program, company outings, and designated volunteering hours with group volunteer events.
Tech Stack
Apache KafkaAWSAzureDockerElasticsearchGitHub ActionsGoGoogle Cloud PlatformGrafanaHelmJavaJenkinsKubernetesLinuxPostgreSQLPrometheusPythonTerraform
Categories
DevOpsSite Reliability
About Sight Machine
Sight Machine provides a manufacturing data platform that unifies plant-floor and enterprise sources to deliver real-time production monitoring, quality analytics, and decision support for industrial manufacturers. It sells subscription software and services to large manufacturers, runs on a standardized infrastructure-as-code deployment model, and is privately held; founded in 2011, it is headquartered in San Francisco with an office in Ann Arbor.