Lifted

Copy of Senior Python Developer (AI Evaluation & Benchmarking)

Lifted
Apply
2 months ago
Remote, PhilippinesSenior

Responsibilities

  • Design and develop coding benchmarks for evaluating frontier AI models.
  • Analyze AI-generated code for correctness, reliability, efficiency, and edge cases.
  • Build and maintain scalable data pipelines for AI evaluation workflows.
  • Create structured programming scenarios testing reasoning, debugging, and code quality.
  • Work with large codebases and multi-language software environments.
  • Write clean, maintainable, and well-tested Python code.
  • Collaborate on improving AI models’ software understanding, generation, and evaluation.

Requirements

  • At least 4 years of professional software engineering experience.
  • Expert-level proficiency in Python.
  • Experience at a high-growth technology company or top-tier software organization.
  • Proficiency in at least one additional programming language such as JavaScript, Go, or C++.
  • Experience with automated testing frameworks such as pytest, Mocha, or JUnit.
  • Understanding of software engineering best practices, debugging, and code quality.
  • Strong analytical and problem-solving skills.
  • Experience with AI/ML evaluation, model benchmarking, or generative AI is preferred.
  • Background in security engineering is preferred.
  • Significant open-source software contributions are preferred.
  • Experience with large-scale distributed systems or enterprise software platforms is preferred.

Benefits

  • Fully remote contract opportunity.
  • Compensation of $80–$100 USD per hour.
  • Expected workload of 10–39 hours per week depending on project needs.
  • Weekly payments for approved work completed during the previous week.
  • Work volume may fluctuate throughout the engagement.
  • Hiring includes a proposal, qualification form, client evaluation, technical interview, and onboarding process.

Tech Stack

C++GoJavaScriptJUnitMochapytestPython
Lifted

About Lifted

201-500 employees
Contact me