2 months ago
Base Salary
$176k - $220k/yr
Responsibilities
- Build and operate shared infrastructure for production ML and AI, including data pipelines, feature stores, training, and model serving.
- Develop and scale an LLM platform with provider integrations, orchestration, observability, and controls for cost, latency, and reliability.
- Build LLM evaluation harnesses, benchmarks, and quality measurement pipelines.
- Support fine-tuning, reinforcement-learning, and related post-training workflows and data infrastructure.
- Optimize inference infrastructure for open and fine-tuned models through GPU serving, batching, and autoscaling.
- Partner with AI, Data Science, and Product teams to productionize models and establish ML infrastructure best practices.
- Improve the reliability, scalability, and developer experience of the ML platform.
Requirements
- At least 5 years of production software engineering experience using Python, Go, TypeScript, or similar languages.
- Experience building and operating cloud infrastructure on AWS, GCP, or similar platforms.
- Strong experience with Kubernetes, Docker, Terraform, and operating production services.
- Hands-on experience building ML infrastructure such as model serving, training pipelines, feature stores, embeddings, or ML observability.
- Experience with modern data platforms including BigQuery, Airflow, Spark, Beam/Dataflow, or streaming pipelines.
- Practical experience building production systems with LLMs or generative AI, including orchestration, provider APIs, observability, and performance optimization.
- Strong systems design skills, sound engineering judgment, and the ability to work in ambiguous, fast-moving environments.
- Extra credit for experience with Ray, Anyscale, KubeRay, Ray Serve, vLLM, Triton, PyTorch, or GPU-backed inference and training.
- Extra credit for designing LLM evaluation frameworks, benchmarking systems, or quality regression testing.
- Extra credit for experience with Vertex AI, Bigtable, Redis, or feature platform infrastructure.
- Extra credit for post-training techniques such as fine-tuning, RLHF, reinforcement learning, or reward modeling.
- Extra credit for agentic systems, MCP integrations, tool use, memory systems, or voice AI applications.
Benefits
- Equity in a fast-growing company
- 401(k) match, competitive compensation, and financial coaching
- Paid parental leave, fertility benefits, and parental coaching
- Medical, dental, vision, and mental health support
- $500 wellness stipend
- $2,000 learning stipend and ongoing development
- Internet and commuting support
- Free lunch and gym access in the San Francisco office
- Flexible PTO, 15 holidays, and 2 flex days
- Team outings and referral bonuses
- Benefits listed apply to full-time US employees; remote and office arrangements are supported
Tech Stack
Apache AirflowApache BeamApache SparkAWSDockerGoGoogle BigQueryGoogle Cloud PlatformKubernetesPythonPyTorchRedisTerraformTypeScript
Categories
About Handshake
Handshake builds a career network and recruiting platform that connects college students and recent grads with employers, sold as SaaS to universities and subscriptions/solutions to employers. Founded in 2014 and headquartered in San Francisco, it serves 1,600+ educational institutions and over 1 million employers. The company also operates Handshake AI, which partners with frontier AI labs on human data collection and evaluations for model training and post-training workflows.
