6 months ago
Base Salary
$125k - $225k/yr
Responsibilities
- Write maintainable, scalable, and performant backend code primarily in Go, Java, and Python, with opportunities to use TypeScript.
- Build high-volume, highly available analytics systems and scalable backend services for the core platform.
- Design and build APIs for customers’ machine learning and LLM workflows.
- Extend and contribute to open-source OLAP databases and distributed message queue frameworks.
- Develop and integrate collection tools for monitoring ML and LLM pipelines.
- Research and implement visualization and dimensionality-reduction algorithms in distributed environments.
- Collaborate with product, design, and customer engineering teams to expand Arize’s product offerings.
- Contribute to building in-house AI agents.
Requirements
- At least 5 years of experience working with high-performance backend systems.
- Strong experience with Go, Python, TypeScript/Node, Java, or similar server-side programming languages.
- Experience building and operating highly complex SaaS platforms or systems.
- Knowledge of public clouds and container orchestration, including AWS, GCP, Azure, or Kubernetes.
- Interest in the AI and LLM ecosystem and willingness to learn emerging technologies.
- Preferred experience with distributed stream processing technologies such as Kafka or Gazette.
- Preferred experience with OLAP systems and observability tooling such as Prometheus.
- Working knowledge of machine learning or data science is preferred.
- First-hand experience with large language models or developing AI products is preferred.
Benefits
- Medical, dental, and vision coverage.
- 401(k) plan, unlimited paid time off, generous parental leave, and mental health and wellness support.
- Remote-first work arrangement, with optional in-person offices in New York City and the San Francisco Bay Area.
- WFH monthly stipend for coworking spaces for employees outside those office locations.
- Competitive equity package.
Tech Stack
Categories
About Arize AI
Arize AI builds an observability and evaluation platform for machine learning, LLMs, and AI agents, helping teams monitor, trace, and improve models in production. The SaaS product is used by AI/ML engineers and MLOps teams to detect issues, run evaluations, and optimize performance using real production signals. Founded in 2020 and headquartered in San Francisco, the company serves enterprise customers across industries.
