
Reinforcement Learning Environments Engineer - Cybersecurity
Preference Model3 months ago
Responsibilities
- Design and build reinforcement learning environments and reward functions for offensive and defensive cybersecurity tasks across diverse programming languages.
- Create environments covering vulnerability discovery in source code, exploitation, and safe patching.
- Build reverse engineering environments for binaries, bytecode, and obfuscated code.
- Construct verifiable reward signals using fuzzers, sanitizers, symbolic execution, static analyzers, exploit-success checks, and patch-correctness validation.
- Collaborate with the team to develop ideas and tools that improve the environment-building process.
Requirements
- Strong security fundamentals across offensive and defensive security, including the ability to understand vulnerabilities and translate them into reinforcement learning problems.
- Hands-on experience finding, exploiting, or patching real vulnerabilities through CTFs, bug bounty work, security research, red/blue team engagements, or industry security work.
- Proficiency in Python and systems programming, plus working comfort with at least one of C, C++, or Rust and one web or application stack.
- Familiarity with fuzzers, sanitizers, debuggers, and disassemblers.
- Published security research, CVEs, notable bug bounty findings, strong CTF or competitive results, or deep expertise in a security specialization are valued.
- Experience building or contributing to fuzzing infrastructure, vulnerability scanners, automated program analysis tools, machine learning for code or security, complex interactive reinforcement learning environments, agent harnesses, or sandboxed evaluation infrastructure is valued.
- Ability to take ownership, drive solutions end-to-end, meet throughput expectations, respond quickly to feedback, and stay current with security and machine learning developments.
Benefits
- Competitive cash and equity compensation, with compensation stated as above the 90th percentile
- Health, vision, and dental benefits
- 401(k) match
- Visa sponsorship and relocation support available
- Ownership and autonomy in a fast-moving startup environment
- Opportunity to work with top machine learning engineers
About Preference Model
Preference Model builds RL environments that automate ML research and engineering.