Sunset Data Scientist at Sunset focused on evaluation of de-identified enterprise data quality and safety. This hands-on role involves building datasets, experiments, quality measures, and feedback loops to improve models and pipelines.
Responsibilities
Sunset turns sensitive internal enterprise data into de-identified datasets without destroying the structure and meaning that make the data valuable. That creates a difficult measurement problem. A system can improve aggregate F1 while missing a high-risk slice, remove more sensitive information while also destroying useful context, or pass one stage while defects escape somewhere else in the pipeline.
As Sunset's first Data Scientist focused on evaluation, you will establish how we know whether that data is actually getting better. You will build the datasets, experiments, quality measures, and feedback loops that expose hidden failures, accelerate model and pipeline improvement, and give the team confidence in what it delivers.
This is a hands-on, zero-to-one role at the intersection of data science, AI, and a real production system. You will write Python and SQL, construct evaluation corpora, study failure patterns, design comparisons, calibrate human and model-based judgments, and turn the result into a clear decision. The questions are scientifically difficult, but the output must be practical enough to change what the team builds and ships.
Questions You Might Answer
Did a higher NER or entity-resolution score actually reduce sensitive misses across the messages, documents, tables, and providers that matter?
Is a new model finding more sensitive information, or simply removing more of the useful structure our customers need?
Can we trust a golden dataset, a human review process, or an LLM judge enough to use it for a release decision?
Which customer, modality, entity, language, or format slices are hidden by a strong aggregate result?
Where did a quality loss enter between source data, processing, de-identification, review, and delivery?
What is the smallest credible experiment that would tell us whether to ship, revise, or stop a change?
Define what high-quality and safe-to-deliver data mean across de-identification, structure preservation, semantic coherence, and customer utility
Design representative samples and build golden, adversarial, replay, and production-like corpora with explicit provenance, labeling policy, agreement, adjudication, and versioning
Qualification
You understand samplingYou can investigate messyThis Role May Not Be for You IfYou treat labels
Required
You understand sampling, uncertainty, precision, recall, F1, calibration, agreement, class imbalance, distribution shift, and imperfect labels
You can investigate messy, multi-stage data systems and determine where an apparent gain or loss actually came from
You are comfortable writing Python and SQL and building reproducible technical artifacts rather than handing requirements to an engineering team
You can protect the independence of an evaluation while collaborating closely with the people whose work it evaluates
You use AI tools fluently but do not confuse an articulate model output with valid evidence
You communicate uncertainty and difficult findings directly, without hiding behind false precision
This Role May Not Be for You If
You want to optimize models as your primary job rather than determine whether changes actually improve delivered data
You prefer descriptive dashboards that stop short of changing a decision
You treat labels, benchmarks, or model-based judges as ground truth without investigating how they fail
You need a perfectly defined dataset and research plan before you can make progress
You are uncomfortable disagreeing with a technically strong team when the evidence does not support its conclusion