
6/15/2026
What this post added
Introduced an automated data curation pipeline using LLM judge consensus to train Sidekick's models to refuse impossible or ambiguous requests. This involved calibrating LLM judges with a small seed dataset, establishing a strict consensus mechanism for labeling, and defining a taxonomy of refusal categories. The pipeline creates a data flywheel where production traffic from improved models fuels subsequent training runs, leading to significant gains in skill evaluation scores.