Preference Model

Reinforcement Learning Environments Engineer - Cybersecurity

Preference Model
Apply
3 months ago

Responsibilities

  • Design and build reinforcement learning environments and reward functions for offensive and defensive cybersecurity tasks across diverse programming languages.
  • Create environments covering vulnerability discovery in source code, exploitation, and safe patching.
  • Build reverse engineering environments for binaries, bytecode, and obfuscated code.
  • Construct verifiable reward signals using fuzzers, sanitizers, symbolic execution, static analyzers, exploit-success checks, and patch-correctness validation.
  • Collaborate with the team to develop ideas and tools that improve the environment-building process.

Requirements

  • Strong security fundamentals across offensive and defensive security, including the ability to understand vulnerabilities and translate them into reinforcement learning problems.
  • Hands-on experience finding, exploiting, or patching real vulnerabilities through CTFs, bug bounty work, security research, red/blue team engagements, or industry security work.
  • Proficiency in Python and systems programming, plus working comfort with at least one of C, C++, or Rust and one web or application stack.
  • Familiarity with fuzzers, sanitizers, debuggers, and disassemblers.
  • Published security research, CVEs, notable bug bounty findings, strong CTF or competitive results, or deep expertise in a security specialization are valued.
  • Experience building or contributing to fuzzing infrastructure, vulnerability scanners, automated program analysis tools, machine learning for code or security, complex interactive reinforcement learning environments, agent harnesses, or sandboxed evaluation infrastructure is valued.
  • Ability to take ownership, drive solutions end-to-end, meet throughput expectations, respond quickly to feedback, and stay current with security and machine learning developments.

Benefits

  • Competitive cash and equity compensation, with compensation stated as above the 90th percentile
  • Health, vision, and dental benefits
  • 401(k) match
  • Visa sponsorship and relocation support available
  • Ownership and autonomy in a fast-moving startup environment
  • Opportunity to work with top machine learning engineers

Tech Stack

Categories

Preference Model

About Preference Model

11-50 employees

Preference Model builds RL environments that automate ML research and engineering.