17 hours ago
Base Salary
$130k - $225k/yr
Responsibilities
- Build systems to detect and transform PII, quasi-identifiers, credentials, and other sensitive information.
- Develop and benchmark rules-based, statistical, classifier-based, and LLM-based detection approaches.
- Create production anonymization pipelines for processing, training, evaluation, and synthetic data workflows.
- Develop privacy-risk and retained-utility evaluation frameworks, including leakage tests and adversarial re-identification attempts.
- Design systems that handle new sources, schema drift, unusual formats, and sensitive information in unexpected fields.
- Partner with engineering, research, operations, and customers to translate privacy requirements into practical safeguards.
Requirements
- At least 2 years of experience building production data or ML systems in Python, with strong Python proficiency.
- Hands-on experience with PII detection, removal, or anonymization and utility-preserving transformations.
- Experience with information extraction, named-entity recognition, classification, or related methods for finding rare or sensitive content.
- Ability to build end-to-end data pipelines and compare approaches using recall, precision, latency, cost, and downstream utility.
- Understanding of redaction, masking, pseudonymization, anonymization, and synthetic data generation.
- Experience handling schema drift and edge cases.
- Experience with differential privacy, k-anonymity, secure aggregation, or format-preserving encryption is valuable.
- Experience with low-latency or high-throughput ML inference and data processing is beneficial.
Benefits
- Salary range of $130,000 to $225,000 annually.
- Visa sponsorship is available.
- On-site role in San Francisco, California, United States.
