
Member of Technical Staff (EMEA)
Fireworks AI3 hours ago
London, United KingdomMid Level / Senior
H1B Sponsor
Responsibilities
- Design and ship product capabilities end to end, from ambiguous needs through production deployment.
- Build inference and serving capabilities including routing, model controls, observability, and deployment functionality.
- Develop tooling for LLM fine-tuning and post-training, including data pipelines and evaluation.
- Translate EMEA-specific requirements into product and coordinate with the core platform team.
- Set the engineering quality bar for the regional team through high-quality production code.
- Use coding agents and AI tooling as a core part of development with sound judgment about their output.
- Optionally participate in customer deployments and conversations.
Requirements
- Demonstrate first-principles thinking and the ability to reason through novel problems without a reference implementation.
- Have excellent software engineering depth in backend systems and infrastructure.
- Have a track record of shipping and owning production systems.
- Be highly proficient in Python and at least one systems language.
- Be fluent in AI-assisted and agentic engineering, including coding agents and AI tooling.
- Have a genuine interest in inference and AI infrastructure.
Benefits
- Work on challenging AI infrastructure problems including low-latency inference and scalable model serving.
- Build technology that shapes how businesses and developers use AI globally.
- Have significant ownership and impact in a fast-growing, low-bureaucracy environment.
- Collaborate with experienced engineers and AI researchers.
- Fireworks AI is an equal-opportunity employer committed to an inclusive workplace.
Tech Stack
Categories
About Fireworks AI
Fireworks is the fastest way to build, tune, and scale AI on open models. Ship production-ready AI in seconds on our globally distributed cloud infrastructure, optimized for your use case. Fireworks powers production workloads at companies like Uber, Doordash, Notion, and Cursor—delivering 15× faster speed, 4× lower latency, and 4× more concurrency than closed models.