3 hours ago
Atlanta, GA, USASenior
Base Salary
$116k - $175k/yr
Responsibilities
- Partner with Engineering, Operations, and Product teams to design, deliver, and maintain highly available, performant application platforms.
- Build application observability and platform monitoring tools to improve customer experience.
- Automate operational processes, eliminate toil, tune alerts, and improve code where needed.
- Define, implement, and maintain SLIs and SLOs aligned with customer experience.
- Instrument latency, error rate, and availability measurements across critical services.
- Manage error budgets and balance system reliability with product feature velocity.
- Improve alert quality by reducing noise and emphasizing actionable, high-signal alerts.
- Embed with product teams to review architectures and identify reliability risks early.
- Build Python, Bash, Java, or Ruby scripts for operational automation and incident response.
- Build or operate AI-assisted incident response systems for root-cause analysis, log summarization, and anomaly triage.
- Develop or integrate LLM-based tools to reduce MTTR and improve alert quality.
- Apply machine learning techniques to anomaly detection, capacity prediction, and failure pattern analysis.
- Deploy AI systems in production and use vector databases, embeddings, or RAG architectures for operational intelligence.
- Share technical knowledge and findings with engineering teams, technical leadership, and senior management.
Requirements
- Bachelor's degree in computer science, engineering, or a related technical or business field.
- At least 4 years of application development experience with Java or an equivalent language.
- Experience with the Spring environment and cloud-based infrastructure such as Azure, AWS, or GCP.
- Understanding of database, network, CPU, JVM, memory, thread-management, and query-performance factors affecting applications.
- Knowledge of centralized logging, metrics dashboards, and alerting tools.
- Awareness of SQL and NoSQL databases.
- Hands-on experience with Datadog, Prometheus, Grafana, or similar observability tools.
- Knowledge of CI/CD pipelines and infrastructure-as-code using Terraform, Helm, Jenkins, or GitLab.
- Experience with AI-assisted incident response, LLM-based tools, machine learning for operational analysis, and production AI deployments.
- Knowledge of vector databases, embeddings, RAG architectures, prompt engineering, and evaluation of LLM outputs in reliability workflows.
- Experience with Kubernetes and container orchestration, including EKS, AKS, or GKE.
- Experience with distributed systems at scale and familiarity with service meshes and microservices architectures.
- Nice-to-have experience with Gremlin or Chaos Monkey, high-traffic product-facing services, and incident management platforms such as PagerDuty and Datadog.
Benefits
- Comprehensive healthcare coverage.
- Flexible paid time off.
- Equity RSUs and annual performance bonus opportunities.
- Retirement account support.
- 14+ weeks of paid parental leave.
- Career development opportunities.
- Company-paid privacy certification exam fees.
- Office-first work culture encouraging three days per week in the office for most roles.
- Additional benefits vary by country.
Tech Stack
AWSAzureBashDatadogGoogle Cloud PlatformGrafanaHelmJavaJenkinsKubernetesPrometheusPythonRubySQLTerraform
Categories
Site Reliability
About OneTrust
OneTrust, the AI-Ready Governance Platform™, enables innovation through the responsible use of data and AI. Trusted by over half of the Fortune 500, we help businesses govern well and move fast, turning responsible data use into a catalyst for growth. To learn more, visit www.onetrust.com