The user's question is deceptively simple: given a pile of thumbs up and thumbs down, what do you actually do? Most practitioners default to the obvious move, calculating approval percentages and training a reward model for RLHF. That instinct is not wrong, but it undersells the problem. A binary label on a single response tells you what the user did not like, but it rarely tells you why, and without that context, you are optimizing for a proxy that can drift from user intent. The dataset is not a limitation to work around; it is the raw material for a more disciplined evaluation strategy.
What stands out here is the constraint that you cannot generate new responses. That forces a focus on offline evaluation, which is actually a gift. It pushes you to think carefully about what your existing labels can and cannot support. A simple approval rate is noisy, especially if the distribution of prompts is skewed or if some responses are easier to satisfy than others. You can stratify by prompt type, by response length, by domain, and look for patterns in where the thumbs down cluster. That is not a substitute for a reward model, but it is a necessary diagnostic layer. The reward model itself should be trained with the understanding that you are predicting a preference, not a quality score. That distinction matters, and it changes how you frame the loss and what you expect from the model.
The deeper issue is that RLHF on a fixed, small preference dataset is brittle. You risk overfitting to the exact distribution of the thumbs down you already have, which means you improve on past failures but stay blind to new ones. The better path is to treat the preference labels as a seed, not a final answer. Use them to build a reward model that can generalize, but also consider simpler interventions: supervised fine-tuning on the high-scoring responses, or even just filtering and rebalancing your training data. Those are less glamorous than a full RLHF loop, but they are often more robust when you cannot iterate online.
The literature does support this. Preference optimization methods, whether you call it RLHF or direct preference optimization, all assume you have a reliable signal. The signal here is user clicks, which are honest but incomplete. So the practical recommendation is not to chase the fanciest algorithm. It is to build a tight loop between evaluation and fine-tuning, using the labels to identify failure modes, then addressing those specific modes with targeted training. That is how you turn a pile of thumbs into a clear path forward, not by asking what the best method is, but by asking what your data is already telling you.