
Callosum
Open Positions at Callosum
13 open positions
Build the foundational runtime for long-running AI agents operating across distributed infrastructure, with durable execution, secure tool use, scheduling, and deep observability. This high-leverage role combines distributed systems, framework design, and runtime engineering from first principles.
Callosum is seeking an ML Research Engineer to investigate why agentic AI systems fail and build evidence-driven methods for improving their intelligence, cost, and reliability. The role combines independent research, experimental engineering, agent evaluation infrastructure, and publication in a London-based team working across heterogeneous AI hardware.
Own and advance Callosum’s LLM-guided program evolution and optimisation system, enabling automated search across models, software, hardware, and workloads. This research-heavy engineering role combines evolutionary optimisation, search strategy design, and reliable internal infrastructure.
Own and build Callosum’s rigorous benchmarking system for agentic and algorithmic LLM approaches, using reproducible execution-based evaluation rather than self-reported scores. The role combines research-grade evaluation with hands-on engineering and directly informs product decisions, customer proof points, and published benchmarks.
Own the communication systems that let heterogeneous AI accelerators operate as one efficient system. You’ll improve data movement, networking, memory transfer, and interconnect performance across emerging hardware platforms.
Own the operational health and reliability of Callosum’s customer-facing AI infrastructure platform as it scales across heterogeneous compute backends. You’ll define the reliability practice, lead incident response, and shape production operations for a rapidly growing AI infrastructure company.
Build and deploy optimized AI workflows for customer production systems as an early Applied AI engineer at Callosum. You’ll lead technical evaluations, integrations, and engagements while shaping the team’s methodology and product roadmap.
Own inference performance and deployment for heterogeneous AI infrastructure, running production-equivalent experiments across cloud and on-prem hardware. You’ll build reproducible deployment patterns, benchmarking systems, and orchestration improvements that guide the stack toward production readiness.
Build the next generation of heterogeneous AI inference engines by extending systems such as vLLM and SGLang for hardware-aware scheduling, memory management, execution, and serving. You’ll design production-scale parallelism and disaggregation strategies across diverse accelerators.
Build the cloud and cluster orchestration layer that makes heterogeneous AI accelerators deployable, schedulable, and efficiently utilized across providers and regions. This hands-on role combines Kubernetes internals, multi-cloud infrastructure, and accelerator-aware resource management.
Unlock 3 More Matching Jobs
Create an account to view all jobs matching your criteria.
First to know
Discover the latest jobs before everyone else
Zero Spam
No ghost jobs, reposts, or sponsored listings
AI-powered filters
Find the most relevant jobs for you