Together AI

Senior Software Engineer, Observability

Together AI
Apply
3 months ago

Base Salary

$200k - $280k/yr

Responsibilities

  • Design and implement a scalable observability platform for metrics, logs, traces, telemetry pipelines, and log aggregation.
  • Develop automated monitoring, alerting, anomaly detection, SLIs/SLOs, runbooks, and predictive analytics for critical services.
  • Build custom observability tools and infrastructure as code using Go, Python, Terraform, Ansible, and Helm.
  • Improve distributed tracing and application monitoring in collaboration with engineering teams.
  • Lead incident response, conduct post-mortem analysis, and define observability best practices.

Requirements

  • Expertise with observability platforms including Prometheus, Grafana, ClickStack, and OpenTelemetry, as well as cloud-native monitoring across AWS, GCP, and Azure.
  • Strong programming skills in Go, Python, or similar languages and proficiency with Terraform, Ansible, and Helm.
  • Experience designing, operating, and scaling distributed systems and high-volume, real-time data pipelines.
  • Deep understanding of Docker and Kubernetes.
  • Knowledge of microservices architecture, service mesh technologies, CI/CD pipelines, and GitOps workflows.
  • Expertise managing PostgreSQL, MongoDB, Redis, and time-series databases with high-cardinality data.
  • Preferred experience includes monitoring AI/ML infrastructure and GPU clusters, low-latency systems monitoring, chaos engineering, reliability testing, open-source observability contributions, and security or compliance monitoring.

Benefits

  • Full-time position with remote-work flexibility
  • Startup equity
  • Health insurance
  • Other benefits

Tech Stack

AnsibleAWSAzureClickHouseDockerGoGoogle Cloud PlatformGrafanaHelmKubernetesMongoDBPostgreSQLPrometheusPythonRedisTerraform

Categories

DevOpsSite Reliability
Together AI

About Together AI

201-500 employees

Together AI builds an AI-native cloud platform for developers, offering high-performance inference, fine-tuning/model shaping, and large-scale pre-training on on-demand GPU clusters with APIs and managed services. It emphasizes open-source models that teams can run and adapt, and also provides infrastructure for decentralized and scalable workloads. Founded in 2022 and headquartered in San Francisco, it is privately held and reports notable customers including Cursor, ElevenLabs, Salesforce, and Zoom.

Contact me