5 months ago
Responsibilities
- Implement and scale training pipelines for large transformer and LLM models from data ingestion and preprocessing through distributed training and evaluation.
- Build and optimize low-latency, highly reliable inference services, including autoscaling, routing, and fallbacks.
- Tune GPU kernels, improve utilization, and identify bottlenecks across the training and inference stack.
- Collaborate with ML scientists to implement advanced training and inference methods and bring them to production.
- Participate in hiring, mentoring, and developing other engineers.
- Improve technical standards, reliability, and operational excellence across the AI platform.
Requirements
- 5+ years of software engineering experience.
- A degree in Computer Science, Computer Engineering, or a related field, or equivalent experience with very strong fundamentals.
- Hands-on experience with model training, especially transformers and LLMs; model inference at scale; or low-level GPU work such as CUDA or Triton kernels.
- Experience working in production environments at meaningful traffic, data, or organizational scale.
- Deep knowledge of at least one programming language, such as Python, Ruby, Java, or Go.
- Clear communication, strong technical fundamentals, collaborative working style, and willingness to invest in professional development.
- Preferred experience running training or inference workloads on Kubernetes.
- Preferred experience with AWS or other major cloud providers.
- Preferred production experience with Python in ML or infrastructure contexts.
- Experience at an AI-native company that trains or runs inference for its own models is a bonus.
- Personal projects, open-source contributions, meetups, or published technical content are valued.
Benefits
- Competitive salary and equity are offered, with regular compensation reviews; compensation figures are not stated.
- Lunch is provided every weekday, along with snacks and a stocked kitchen.
- Unlimited access to Claude Code and other AI tools is provided.
- Pension scheme with matching up to 4%.
- Life assurance and comprehensive health and dental insurance for employees and dependents.
- Flexible paid time off policy.
- Paid maternity leave and six weeks of paternity leave.
- Cycle-to-Work Scheme and secure bike storage.
- MacBooks are standard, with Windows available for certain roles.
- Hybrid working policy requiring employees to work in the office at least three days per week.
Tech Stack
Categories
About Intercom
We’re Intercom — the AI customer service company helping businesses deliver incredible customer experiences at scale. Our platform combines Fin, the #1 AI Agent for customer service, with our next-generation Helpdesk, a modern workspace that gives support teams the power, speed and intelligence they need.