over 1 year ago
Base Salary
$230k - $325k/yr
Responsibilities
- Design and operate data storage pipelines at scale.
- Build search and discovery services using systems such as Spark, Databricks, index layers, and metadata catalogs.
- Automate secure data transfers with encryption, checksumming, and auditing for reviewer exports.
- Establish locked-down compute environments that balance usability with security controls.
- Instrument monitoring and KPIs for data holds and productions.
- Collaborate across teams to codify SOPs, threat models, and chain-of-custody documentation.
Requirements
- Hands-on experience building or operating large-scale data-lake or backup systems.
- Experience with at least one of Azure, AWS, or GCP.
- Knowledge of Terraform or Pulumi and CI/CD.
- Ability to turn ad-hoc legal requests into repeatable pipelines.
- Experience with discovery workflows such as legal holds, enterprise document collections, or secure review, or willingness to develop that expertise quickly.
- Ability to communicate technical concepts clearly to Legal, Engineering, and other teams.
- Experience shipping secure solutions that balance speed, cost, and evidentiary defensibility.
- Strong documentation and cross-disciplinary communication skills, including the ability to work under tight deadlines.
Benefits
- The position is located in San Francisco.
- Relocation assistance is available.
Tech Stack
Categories
Data Engineering
About OpenAI
OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.
