
Staff Site Reliability Engineer
Sight Machine21 days ago
Remote, United StatesStaff+
Responsibilities
- Drive reliability, automation, scalability, and technical direction across Sight Machine’s cloud infrastructure.
- Design, build, and operate infrastructure for agentic AI workloads, LLM gateways, agent orchestration, and non-deterministic services.
- Evolve SLO, error-budget, incident-response, postmortem, and reliability-review practices.
- Troubleshoot complex issues spanning CI/CD, container orchestration, networking, operating systems, cloud resources, databases, and AI/LLM orchestration.
- Build monitoring, alerting, observability, operational runbooks, automation, internal platforms, and developer tooling.
- Participate in on-call coverage, improve escalation paths, and reduce avoidable pages through automation.
- Mentor senior and mid-level engineers and influence architecture decisions across teams without formal authority.
- Lead self-directed cross-team initiatives that improve stability, reliability, and availability.
Requirements
- 10+ years of experience with Kubernetes and Docker in at least one of Azure, GCP, or AWS, including production-scale multi-tenant or multi-cluster environments.
- 10+ years of coding experience with Python, Go, Java, or similar, including building tools and platforms used by other engineers.
- 10+ years of experience with infrastructure as code and CI/CD tooling such as Terraform or OpenTofu, FluxCD or similar GitOps tooling, and Jenkins or GitHub Actions.
- Demonstrated production experience designing, building, or operating agentic AI or LLM-based systems.
- Strong Linux and TCP/IP and application-layer networking fundamentals.
- Practical experience integrating or operating production LLM gateways, API-based orchestration, or agent frameworks.
- Operational experience with monitoring and alerting systems such as Prometheus, Grafana, Loki, Sentry, and Signoz or equivalents.
- Deep understanding of cloud performance and the ability to diagnose complex bottlenecks.
- Experience authoring technical documentation, including design documents, ADRs, and runbooks.
- Demonstrated mentorship experience without formal management authority and the ability to influence architecture decisions across teams.
- Strong hands-on judgment under pressure, including balancing technical risk with customer impact.
- Preferred experience with Kubernetes, FluxCD, Terraform, Helm Charts, Prometheus, Elasticsearch, Python, Java, Kafka, Postgres, and Jenkins.
- Interest or experience in industrial IoT, analytics, or manufacturing is a plus.
Benefits
- Competitive salary and stock options.
- Health care coverage, life insurance, health savings account, and flexible spending account, including spouse and children.
- Flexible vacation policy.
- Adaptable working schedule and environment.
- Hybrid work flexibility, with the ideal candidate near the San Francisco or Ann Arbor office; exceptional fully remote candidates may be considered.
- Catered lunches, snacks, and beverages.
- Commuter Savings Program.
- Company outings.
- Designated volunteering hours and group volunteer events.
Tech Stack
Apache KafkaAWSAzureDockerElasticsearchGitHub ActionsGoGoogle Cloud PlatformGrafanaJavaJenkinsKubernetesLinuxPostgreSQLPrometheusPythonTerraform
Categories
DevOpsSite Reliability
About Sight Machine
Sight Machine provides a manufacturing data platform that unifies plant-floor and enterprise sources to deliver real-time production monitoring, quality analytics, and decision support for industrial manufacturers. It sells subscription software and services to large manufacturers, runs on a standardized infrastructure-as-code deployment model, and is privately held; founded in 2011, it is headquartered in San Francisco with an office in Ann Arbor.