5 months ago
Responsibilities
- Design, build, deploy, operate, and maintain platform backend services and data pipelines for AI products.
- Build ML experiment tracking, artifact management, and automated training and evaluation systems.
- Engineer highly available, fault-tolerant systems with deep observability for strict uptime and latency requirements.
- Design modular architectures and clean APIs while optimizing latency, throughput, and cloud compute costs.
- Partner with engineering, product, and ML research teams to build shared platform capabilities and reduce AI product delivery bottlenecks.
Requirements
- At least 3 years of experience building and scaling production backend services or platforms.
- Experience with languages common in the AI and data ecosystem, such as Python, Go, Rust, or C++.
- Deep understanding of consistency, availability, distributed failure modes, and idempotency.
- Ability to build intuitive, well-documented, maintainable APIs and shared platforms.
- Experience owning services with uptime expectations, designing for observability, and participating in on-call responsibilities.
- Ability to turn complex, open-ended problems into clear technical designs and production-ready systems.
- Bonus: experience with ML experiment tracking, artifact management, model registries, or ML orchestrators.
- Bonus: experience with high-throughput data pipelines, resilient asynchronous workflows, streaming platforms, or distributed compute frameworks.
- Bonus: experience with SOC2, FedRAMP, highly regulated, air-gapped, or on-premise environments.
Benefits
- Competitive salary plus equity
- Daily lunches
- Commuter benefits
- 401(k)
- Medical, Dental, and Vision
- Unlimited PTO
Tech Stack
Categories
About Brain Co.
We're building an AI platform and applications for the world's most important institutions. Learn more at https://brain.co/
