In research published October 6, OpenAI described LASER, a pipeline that uses inexpensive classifiers to select informative conversations, then reasoning models to label them. It repeats that process to assemble diverse evaluation examples near difficult policy boundaries. The pipeline uses synthetic and de-identified conversations, rather than accessing raw user data. OpenAI reports that, when finding comparable numbers of disallowed examples, the approach can need roughly 10,000 times less grader compute than random sampling. This is a task-specific research result, not a promise of the same saving for every evaluation pipeline. For builders, the useful principle is to spend expensive evaluation effort on examples that are uncertain and diverse. A practical learning exercise is to compare a randomly sampled test set with one deliberately covering borderline cases, keeping the evaluation criteria fixed. That suggested exercise is an editorial application of the research, rather than a released OpenAI product.
OpenAI research · Announced 6 October 2026