6 months ago
Amsterdam, NetherlandsSenior
Responsibilities
- Improve the YTsaurus core for the TractoAI platform
- Design new system features and microservices
- Investigate performance issues in user workloads
- Collaborate with the SRE team to improve system stability
- Support S3 as storage alongside physical disks
- Implement fair disk-throughput distribution across concurrent user computations
- Enable high-end GPUs and InfiniBand connectors inside user-job VMs and collect utilization statistics
- Improve the TractoAI scheduler’s MapReduce input-data splitting
- Optimize TractoAI I/O engines for HDDs and NVMe SSDs
Requirements
- 5+ years of experience as a software engineer
- Experience with C++ and concurrency
- Ability to take features from ideation through client adoption
- Willingness to investigate unfamiliar areas, including HPC or GPU computing
- Willingness to occasionally write code in Go and Python
- Experience with distributed systems design is a bonus
- Understanding of database internals is a bonus
- Operating-system internals knowledge is a bonus
- Performance engineering experience is a bonus
- Experience maintaining large-scale stateful systems is a bonus
- Participation in coding interviews is required
Benefits
- Flexible working arrangements
- Competitive salary and comprehensive benefits package
- Opportunities for professional growth
- Dynamic and collaborative work environment
About Nebius
The Nebius AI Cloud brings powerful full-stack infrastructure for AI developers and practitioners across startups, enterprises and science institutes to build and deploy generative AI applications and rapidly deliver scientific breakthroughs by training and running ML models within a secure, high-performance, and cost-optimized cloud environment.