2 months ago
Base Salary
$116k - $175k/yr
Responsibilities
- Partner with engineering, operations, and product teams to design, deliver, and maintain highly available and performant application platforms.
- Build application observability and platform monitoring tools and improve customer experience.
- Automate operational processes, eliminate toil, tune alerts, and improve code where needed.
- Define, implement, and maintain SLIs and SLOs aligned with customer experience.
- Instrument latency, error rate, and availability SLIs across critical services and manage error budgets.
- Improve alert quality by reducing noise and focusing on actionable, high-signal alerts.
- Embed with product teams to review architectures and identify reliability risks early.
- Build Python, Bash, Java, or Ruby scripts for operational automation and incident response.
- Build and operate AI-assisted incident-response systems for root-cause analysis, log summarization, and anomaly triage.
- Develop or integrate LLM-based tools, apply machine learning for anomaly detection and capacity prediction, and share findings with technical leadership and senior management.
- Evaluate new tools and techniques, collaborate across functional groups to resolve issues, and share knowledge with the engineering organization.
Requirements
- Bachelor's degree in computer science, engineering, or a related technical or business field.
- At least four years of application development experience with Java or an equivalent language.
- Experience with the Spring environment and cloud-based infrastructure such as Azure, AWS, or GCP.
- Understanding of application performance factors including database and network performance, CPU utilization, JVM tuning, memory analysis, thread management, and query performance.
- Knowledge of centralized logging, metrics dashboards, alerting, databases, and observability tools such as Datadog, Prometheus, and Grafana.
- Knowledge of CI/CD pipelines and infrastructure-as-code tools including Terraform, Helm, Jenkins, and GitLab.
- Experience building or operating AI-assisted incident-response systems, developing or integrating LLM-based tools, applying machine learning techniques, deploying AI systems in production, and using vector databases, embeddings, or RAG architectures.
- Knowledge of prompt engineering and evaluation of LLM outputs in reliability workflows.
- Experience with Kubernetes, container orchestration, distributed systems at scale, service meshes, and microservices architectures.
- Nice-to-have experience with chaos engineering tools such as Gremlin or Chaos Monkey, high-traffic product-facing services, incident management platforms such as PagerDuty, and Datadog monitoring.
Benefits
- Office-first culture with three days per week in the office for most roles
- Comprehensive healthcare coverage
- Flexible paid time off
- Equity RSUs
- Annual performance bonus opportunities
- Retirement account support
- 14+ weeks of paid parental leave
- Career development opportunities
- Company-paid privacy certification exam fees
Tech Stack
AWSAzureBashDatadogGoogle Cloud PlatformGrafanaHelmJavaJenkinsKubernetesPrometheusPythonRubySQLTerraform
Categories
Site Reliability
About OneTrust
OneTrust builds an enterprise data governance and compliance platform used to manage privacy, consent, and AI governance across global operations. Its products span privacy impact assessment and data‑mapping automation, cookie and consent management, vendor risk, subject access requests, incident management, and policy enforcement, sold as subscription software. Founded in 2016 and headquartered in Atlanta, the privately held company serves thousands of organizations worldwide across regulated industries.