4 days ago
New York, NY, USAMid Level
Responsibilities
- Own and improve NER, entity resolution, structured or tabular detection, document understanding, semantic review, and related de-identification systems.
- Turn model failures and capability ceilings into a prioritized improvement roadmap.
- Design active-learning loops using model sweeps, LLM-assisted review, clustering, and uncertainty signals to select examples for labeling.
- Build representative datasets, benchmarks, evaluation environments, seeded failure modes, and programmatic verifiers.
- Design experiments, tune thresholds, analyze precision-recall and utility tradeoffs, and assess whether changes generalize.
- Choose and combine deterministic rules, classical ML, fine-tuning, embeddings, multimodal models, and LLM-based approaches.
- Productionize improvements with reproducible artifacts, evaluation evidence, runtime instrumentation, and safe rollout.
- Optimize inference cost, latency, and throughput without masking quality or high-risk recall regressions.
- Build reliable model- or agent-based harnesses with bounded behavior and explicit output verification when needed.
- Partner with Applied Science, Data and Product Engineering, Security, and Quality on measurement, pipelines, review systems, and acceptable risk.
- Use AI engineering tools to accelerate research, implementation, error analysis, and evaluation while verifying their output.
Requirements
- 3+ years of professional machine learning or software engineering experience, including improving models in production.
- Strong Python and software engineering skills, with the ability to work in data pipelines and production systems.
- Experience moving model quality through error analysis, data work, experimentation, implementation, deployment, and iteration.
- Strong understanding of precision, recall, F1, calibration, thresholding, class imbalance, imperfect labels, distribution shift, and representative evaluation.
- Ability to combine deterministic, statistical, neural, and LLM-based approaches based on the problem.
- Ability to communicate uncertainty and tradeoffs clearly to scientists, engineers, and delivery or risk decision-makers.
- Startup experience and comfort with broad ownership and evolving or incomplete requirements are valued.
- Fluency with modern AI engineering tools and willingness to verify their output are expected.
- Bonus experience includes NER, entity resolution, information extraction, document understanding, multimodal systems, privacy-preserving ML, active learning, uncertainty sampling, weak supervision, and human-in-the-loop review.
- Bonus experience includes fine-tuning or adapting transformer, GLiNER, embedding, vision-language, or specialized models; model optimization; evaluation environments; sensitive enterprise data; and synthetic data generation.
Tech Stack
Categories
About Sunset
We help tech companies shut down. From state withdrawals to liquidations, we save founders thousands of dollars, hundreds of hours, and countless headaches when it comes to winding down their operations.
