2 months ago
Remote, Brazil or São Paulo, BrazilSenior
Responsibilities
- Design and develop coding benchmarks for evaluating frontier AI models.
- Analyze AI-generated code for correctness, reliability, efficiency, and edge cases.
- Build and maintain scalable data pipelines supporting AI evaluation workflows.
- Create programming scenarios that test reasoning, debugging, and code quality.
- Work with large codebases and multi-language software environments.
- Write clean, maintainable, and well-tested Python code.
- Collaborate with teams improving how AI models understand, generate, and evaluate software.
Requirements
- At least 4 years of professional software engineering experience.
- Expert-level proficiency in Python.
- Experience at a high-growth technology company or top-tier software organization.
- Proficiency in at least one additional language such as JavaScript, Go, or C++.
- Experience with CI/CD pipelines and automated testing frameworks such as pytest, Mocha, or JUnit.
- Strong understanding of software engineering best practices, debugging, and code quality.
- Strong analytical and problem-solving skills.
- Preferred experience with AI/ML evaluation, model benchmarking, or generative AI.
- Preferred background in security engineering.
- Significant open-source software contributions are preferred.
- Experience with large-scale distributed systems or enterprise software platforms is preferred.
Benefits
- Fully remote contract opportunity.
- Expected workload is 10–39 hours per week depending on project needs.
- Weekly payments are made for approved work completed during the previous week.
- Work volume may fluctuate throughout the engagement.
- Hiring proceeds through a proposal, qualification form, client contract offer, onboarding, evaluation, and technical interview.
Tech Stack
Categories
BackendData Engineering
