2 months ago
Remote, Canada or Ottawa, CanadaSenior
Responsibilities
- Design and develop coding benchmarks for evaluating frontier AI models.
- Analyze AI-generated code for correctness, reliability, efficiency, and edge cases.
- Build and maintain scalable data pipelines supporting AI evaluation workflows.
- Create programming scenarios testing reasoning, debugging, and code quality.
- Work across large codebases and multilingual software environments.
- Write clean, maintainable, well-tested Python code and collaborate on AI software evaluation improvements.
Requirements
- At least 4 years of professional software engineering experience.
- Expert-level proficiency in Python.
- Experience at a high-growth technology company or top-tier software organization.
- Proficiency in at least one additional language such as JavaScript, Go, or C++.
- Experience with automated testing frameworks such as pytest, Mocha, or JUnit.
- Strong understanding of software engineering best practices, debugging, and code quality.
- Strong analytical and problem-solving skills.
- Preferred: experience with AI/ML evaluation, model benchmarking, generative AI, security engineering, open-source contributions, distributed systems, or enterprise software platforms.
Benefits
- Fully remote contract opportunity.
- Flexible workload of 10–39 hours per week depending on project needs.
- Weekly payments for approved work completed during the previous week.
- Work volume may fluctuate throughout the engagement.
- Applicants complete a proposal and qualification form, followed by an evaluation and technical interview for qualified candidates.
