5 days ago
Responsibilities
- Build systems that detect sensitive information and apply appropriate transformations based on data type and downstream use case.
- Develop and benchmark rules, statistical models, classifiers, and LLM-based methods for sensitive-content detection.
- Build production anonymization pipelines for processing, training, evaluation, and synthetic data workflows.
- Create evaluation frameworks for privacy risk and retained data utility, including leakage tests and adversarial re-identification attempts.
- Design systems robust to new data sources, schema drift, unusual formats, and sensitive information in unexpected fields.
- Collaborate with engineering, research, operations, and customers to translate privacy requirements into technical policies and safeguards.
Requirements
- At least 2 years of experience building reliable production data or machine learning systems in Python.
- Hands-on experience with information extraction, named-entity recognition, classification, or related sensitive-content detection methods.
- Experience building end-to-end data processing pipelines without a fully prescribed roadmap.
- Ability to compare approaches across recall, precision, latency, cost, and downstream data utility.
- Understanding of redaction, masking, pseudonymization, anonymization, and synthetic data generation.
- Experience designing systems robust to schema drift, unusual formats, and edge cases.
- Familiarity with differential privacy, k-anonymity, secure aggregation, or format-preserving encryption is preferred.
- Experience with low-latency or high-throughput machine learning inference and data processing systems is preferred.
- Experience working with sensitive data in healthcare, finance, security, or related domains is preferred.
Benefits
- The role is on-site in San Francisco, California, USA.
- Visa sponsorship is available.
