Lifted

Copy of Senior Python Developer (AI Evaluation & Benchmarking)

Lifted
Apply
2 months ago
Remote, Canada or Ottawa, CanadaSenior

Responsibilities

  • Design and develop coding benchmarks for evaluating frontier AI models.
  • Analyze AI-generated code for correctness, reliability, efficiency, and edge cases.
  • Build and maintain scalable data pipelines supporting AI evaluation workflows.
  • Create programming scenarios testing reasoning, debugging, and code quality.
  • Work across large codebases and multilingual software environments.
  • Write clean, maintainable, well-tested Python code and collaborate on AI software evaluation improvements.

Requirements

  • At least 4 years of professional software engineering experience.
  • Expert-level proficiency in Python.
  • Experience at a high-growth technology company or top-tier software organization.
  • Proficiency in at least one additional language such as JavaScript, Go, or C++.
  • Experience with automated testing frameworks such as pytest, Mocha, or JUnit.
  • Strong understanding of software engineering best practices, debugging, and code quality.
  • Strong analytical and problem-solving skills.
  • Preferred: experience with AI/ML evaluation, model benchmarking, generative AI, security engineering, open-source contributions, distributed systems, or enterprise software platforms.

Benefits

  • Fully remote contract opportunity.
  • Flexible workload of 10–39 hours per week depending on project needs.
  • Weekly payments for approved work completed during the previous week.
  • Work volume may fluctuate throughout the engagement.
  • Applicants complete a proposal and qualification form, followed by an evaluation and technical interview for qualified candidates.

Tech Stack

C++GoJavaScriptJUnitMochapytestPython

Categories

Lifted

About Lifted

201-500 employees
Contact me