Together AI

Staff Platform Engineer, Service Infrastructure

Together AI
Apply
3 months ago

Base Salary

$240k - $280k/yr

Responsibilities

  • Own the technical direction for service infrastructure across Product Foundations, including Kubernetes, AWS, Terraform, CDNs, load balancers, DNS, IAM, and service networking.
  • Improve reliability, operability, deployment safety, infrastructure consistency, and production readiness of existing Product Foundations services.
  • Partner with API Platform and UI Platform on networking, DNS, CDN, load balancing, delivery, and gateway patterns.
  • Coordinate with Infrastructure, Networking, and Security teams on shared platform standards and frameworks.
  • Drive infrastructure initiatives involving Terraform CI/CD, Kubernetes networking, zero-trust service communication, policy-as-code, and cross-provider networking.
  • Build reusable infrastructure primitives such as Helm charts, Terraform modules, GitHub Actions or GitOps workflows, service scaffolding, runbooks, and documentation.
  • Establish technical standards through design documents, architecture reviews, mentorship, and hands-on implementation across teams, regions, and cloud environments.

Requirements

  • 7+ years of professional experience in platform engineering, service infrastructure, SRE, distributed systems, cloud infrastructure, or related roles.
  • Deep production experience with Kubernetes, including EKS, Helm, ArgoCD or Argo Rollouts, ingress, autoscaling, secrets, service identity, networking, and progressive delivery.
  • Strong Terraform experience covering module design, infrastructure CI/CD, policy enforcement, production applies, and safe self-service workflows.
  • Experience operating CDNs, ALBs or NLBs, DNS, TLS, ingress and egress controls, and traffic management.
  • Proficiency in one or more infrastructure automation languages such as Go, Python, TypeScript, or a similar language.
  • AWS experience, ideally including EKS, IAM, VPC networking, load balancing, Route 53, CloudFront, and ECR.
  • Direct experience with observability systems, metrics, logs, traces, dashboards, alerting, SLOs, and incident response.
  • Ability to lead cross-functional technical initiatives across product engineering, infrastructure, networking, and security teams.
  • Strong written communication skills and experience creating design documents, migration plans, operational guidance, and technical standards.
  • Staff-level judgment, including the ability to define ambiguous problems, make pragmatic tradeoffs, influence without authority, and improve systems and teams.
  • Preferred experience includes internal developer platforms, paved-path service frameworks, service mesh or zero-trust infrastructure, policy-as-code, multi-region or multi-provider networking, and supply-chain security.

Benefits

  • Competitive compensation, startup equity, health insurance, and other benefits are offered.
  • This is a full-time position with a US base salary range of $240,000-$280,000 plus equity and benefits.

Tech Stack

AmbassadorAWSGitHub ActionsGoHelmIstioKubernetesPythonTerraformTypeScript

Categories

Together AI

About Together AI

201-500 employees

Together AI builds an AI-native cloud platform for developers, offering high-performance inference, fine-tuning/model shaping, and large-scale pre-training on on-demand GPU clusters with APIs and managed services. It emphasizes open-source models that teams can run and adapt, and also provides infrastructure for decentralized and scalable workloads. Founded in 2022 and headquartered in San Francisco, it is privately held and reports notable customers including Cursor, ElevenLabs, Salesforce, and Zoom.

Contact me