GrepJob
Brain Corp

Senior Staff Software Engineer, Cloud AI Infrastructure

Brain Corp
Apply
5 days ago
Remote, United StatesStaff+

Base Salary

$214k - $267k/yr

Responsibilities

  • Lead and mentor a team of cloud software engineers while providing technical and career guidance.
  • Define and execute the cloud technical roadmap aligned with business and product goals.
  • Architect, implement, and operate scalable, secure, high-availability GCP systems for ML workloads and large-scale data ingestion.
  • Design and build ML pipelines processing hundreds of thousands of images daily for model iteration and deployment.
  • Optimize GPU resource management, model-serving throughput, latency, autoscaling, and cost efficiency.
  • Build canary and staging environments supporting progressive deployments and system resilience.
  • Collaborate with ML, DevOps, data science, and robotics teams on APIs, data models, and cloud-robot workflows.
  • Implement repeatable Infrastructure-as-Code deployments using Pulumi, Terraform, or equivalent tools.
  • Establish cloud observability systems and promote coding, design, distributed-systems, and cloud-ML best practices.

Requirements

  • Bachelor’s or master’s degree in computer science, software engineering, or a related field.
  • 10+ years of professional software engineering experience, including 3+ years in cloud architecture or large-scale distributed systems.
  • Proven experience designing and operating GCP-based machine-learning systems at scale.
  • Expertise with GCP services including GKE, Dataflow, BigQuery, Cloud Run, Pub/Sub, Vertex AI, and Cloud Storage.
  • Strong proficiency in Go, Python, or TypeScript.
  • Experience with ML pipelines, GPU workload optimization, autoscaling, resource scheduling, high-availability systems, and fault-tolerant distributed systems.
  • Hands-on experience with Docker and Kubernetes, plus Pulumi or Terraform and CI/CD systems such as Jenkins or GitHub Actions.
  • Understanding of cloud security, networking, observability, and secure cloud practices.
  • Preferred experience includes robotics data pipelines, fleet management, IoT-scale ingestion, self-hosted ML inference, Vertex AI, Kubeflow, TensorFlow Serving, event-driven architectures, Kafka, and SOC2/ISO27001-compliant systems.
  • Strong problem-solving, communication, leadership, and technical mentorship skills.

Benefits

  • Remote anywhere in the United States or hybrid in San Diego, with relocation available.
  • Discretionary annual target bonus and stock options.
  • 401(k) plan with match, no waiting period, and immediate vesting.
  • Medical, dental, vision, life, disability, HSA, EAP, legal/identity support, and pet insurance benefits.
  • Medical and dependent-care Flexible Spending Accounts.
  • Flexible vacation, paid sick leave, volunteer time off, 10 paid company holidays, and a winter company shutdown.
  • San Diego office perks include daily on-site lunch, an on-campus gym with pool and tennis courts, and colleague events.
  • Internal continuous learning events and opportunities to share personal interests and hobbies.