2 days ago
Princeton, NJ, USASenior / Staff+
H1B sponsor
Responsibilities
- Architect and build end-to-end hierarchical and collaborative multi-agent systems and agentic workflows.
- Develop agent harnesses, orchestration layers, communication protocols, planning mechanisms, state management, memory systems, custom tools, plugins, and tool libraries.
- Implement context engineering, persistent agent interactions, tool integration, self-correction loops, and domain-specific agent capabilities.
- Deploy, scale, maintain, monitor, and optimize production-grade agentic systems on AWS, GCP, or Azure.
- Integrate and optimize LLMs using techniques such as RAG and PEFT.
- Design RAG systems using vector databases and integrate enterprise APIs, databases, and external services.
- Implement monitoring, tracing, observability, evaluation frameworks, production metrics, cost controls, latency analysis, safety controls, and reliability mechanisms.
Requirements
- 6–8 years of hands-on machine learning and AI engineering experience with a proven record of taking ML systems to production.
- Demonstrated expertise building multi-agent systems and agentic workflows, preferably with LangGraph or CrewAI.
- Expert-level Python proficiency and experience with TensorFlow, PyTorch, Transformers, FastAPI, asynchronous programming, and microservices architecture.
- Hands-on experience with Pinecone, Weaviate, or ChromaDB and scalable RAG systems.
- Experience with LLM application monitoring and observability tools such as LangSmith, Weights & Biases, or custom telemetry solutions.
- Ability to architect and implement complex AI systems from scratch in production environments.
- Production experience with at least one of AWS, GCP, or Azure, including compute, serverless, container orchestration, and managed AI/ML services.
- Strong skills in Terraform or CloudFormation, GitHub Actions or Jenkins, Docker, and Kubernetes.
- Preferred experience includes prompt engineering, SLM fine-tuning, PEFT, SFT, RLHF, model optimization, distributed systems, message queues, event-driven architectures, tool-calling agents, agent evaluations, online evaluations, safety controls, cost and latency management, model routing, and memory/state governance.
Tech Stack
AWSAzureDockerFastAPIGitGitHub ActionsGoogle Cloud PlatformJenkinsKubernetesPythonPyTorchReactTensorFlowTerraform
Categories
About Quantiphi
Quantiphi is an AI-first digital engineering and consulting firm that designs and implements machine learning, data, and cloud solutions for large enterprises across industries. Its teams build production systems—such as generative AI applications, computer vision, and predictive analytics—primarily on Google Cloud and other hyperscalers, delivered as professional services and managed solutions. Founded in 2013 and headquartered in Marlborough, Massachusetts, the company is privately held and a Google Cloud partner.
