5 months ago
Base Salary
$140k - $250k/yr
Responsibilities
- Spearhead development of the next-generation multimodal LLM stack.
- Build conversational AI models that power Bland’s agents and take them from research through production.
- Integrate streaming audio, tool execution, and dynamic context into a unified real-time system.
- Define how agents listen, reason, respond, and act in real time.
- Design and run rapid experiments from dataset creation through conclusions.
- Improve agent behavior, latency, correctness, tool selection, policy handling, and real-time understanding.
- Develop systems serving millions of calls per day.
Requirements
- Experience with LLMs, multimodal models, or speech-language systems.
- Deep understanding of prompting, fine-tuning, and alignment techniques.
- Ability to reason about and design complete systems involving models, tools, prompts, and runtime constraints.
- Ability to rapidly design datasets and experiments that produce actionable conclusions.
- Strong product intuition for natural conversational interactions and the ability to translate modeling ideas into user-facing improvements.
- Ownership of work from research through deployment in ambiguous, fast-moving environments.
- Strong attention to latency, correctness, and real-world system behavior.
- Experience with streaming or real-time inference is a plus.
- Experience with real-time voice systems or conversational AI is a bonus.
- Background in tool-using agents or agent frameworks is a bonus.
- Experience with multimodal datasets containing audio, text, and actions is a bonus.
- Contributions to LLM or speech-related research or open source are a bonus.
Benefits
- Meaningful equity is offered.
- Full healthcare, dental, and vision coverage is provided.
- The role is based in San Francisco or remote within the US.
- An office is available in Jackson Square, San Francisco.
- The role offers high autonomy and high impact.
Categories
AI ResearchML Engineering
About Bland
Bland builds AI phone agents for enterprises, providing APIs and infrastructure to automate inbound and outbound calls, route conversations, and integrate with CRMs and contact-center systems. Its platform combines speech-to-text, large language models, neural audio, and text-to-speech to handle customer interactions end to end. Founded in 2023 and privately held, the company targets telecom and enterprise software use cases such as support, scheduling, and sales operations.
