17 days ago
Yerevan, Armenia +9 moreSenior
Responsibilities
- Design and implement tooling to capture, classify, and analyze errors made by AI coding agents generating Kotlin code.
- Build observability pipelines over agentic traces from JetBrains IDEs, Junie, Claude Code, Cursor, and other coding agents.
- Design, implement, and maintain evaluation pipelines measuring Kotlin code-generation quality, including correctness, idiomaticity, build success, framework usage, and test coverage.
- Build simulation environments for realistic Kotlin developer tasks, including greenfield KMP projects, Gradle dependency management, and Java-to-Kotlin Spring migrations.
- Own evaluation infrastructure covering metrics, experiment tracking, automated regression checks, and reproducible benchmarking.
- Research and implement post-training and context-engineering methods to improve agent and model behavior on Kotlin.
- Run A/B comparisons, benchmark suites, and before-and-after analyses on real codebases.
- Collaborate with Anthropic, OpenAI, and Google to translate Kotlin-specific findings into model improvements.
- Design, build, and maintain open-source benchmarks and datasets for AI coding-agent performance on Kotlin tasks.
- Own projects end to end, from identifying failures in agent traces through evaluation design, experimentation, and shipping improvements.
Requirements
- Hands-on experience building evaluation or analysis pipelines for LLMs or AI coding agents in research or production settings.
- At least three years of strong Python engineering experience in clean, maintainable, data-heavy, or ML-adjacent codebases.
- Experience querying large datasets with SQL or Athena, building data pipelines, and performing statistical analysis of experimental results.
- Ability to own projects end to end and translate real developer-facing agent failures into evaluation and training work.
- Product-aware understanding of how developers use coding agents.
- Familiarity with Kotlin or strong willingness to develop deep Kotlin expertise.
- Preferred experience with LLM post-training, including SFT, RLHF, DPO, or GRPO.
- Preferred experience with PyTorch and LLM training stacks such as TRL, verl, or Megatron.
- Preferred experience developing tool-using agents, multi-step coding workflows, or agentic frameworks.
- Preferred experience with Inspect AI, Promptfoo, LM-evaluation-harness, or custom evaluation pipelines.
- Preferred experience with Weights & Biases, MLflow, Langfuse, or similar experiment-tracking and observability tools.
- Preferred familiarity with Android, Gradle, KMP, Spring, Ktor, open-source projects, benchmarks, or evaluation tools.
Benefits
- Competitive base salary reflecting skills and experience.
- Flexible work location with the option to work from home or the office.
- Remote work from abroad for up to 30 days per year.
- Extra time off, medical insurance allowance, learning and development opportunities, and relocation support.
- Language classes, hot meals or a lunch allowance on workdays, mental health support, and sports benefits.
- Internal company events; benefits may vary by location.
About JetBrains
On a mission to make software development a more productive and enjoyable experience. Make it happen. With Code.
