Train smarter data, not just bigger models, for better AI control

In the pursuit of more ethical AI training, the concept of data curation and targeted replacement emerges as a compelling strategy.

3 min readMachine Learning

The most practical path to safer AI may not lie in post-training fixes but in the data we choose to feed the model from the start. A Reddit user, working with limited resources on WikiText-103, has demonstrated that replacing or removing violent content from a training dataset can nearly eliminate violent output while preserving coherence and cohesion. This is not a theoretical argument. It is a working prototype, constrained only by funding and scale.

For anyone building or deploying AI systems, this shifts the conversation. Post-training methods like RLHF and constitutional AI are valuable, but they operate like behavioral therapy on an already-formed mind. They try to suppress what was learned. The alternative, curating the dataset to exclude undesirable material, treats the problem at its source. The user's two methods are instructive: either replace harmful passages with non-violent alternatives that preserve narrative flow and factual accuracy, or, in embedding-based architectures, swap violent tokens for non-violent ones that minimize Hamming distance. Neither approach requires exotic infrastructure. They require intentionality in data preparation.

The open questions are significant and deserve rigorous study. How much does removing deception or violence from training data affect general reasoning, or the model's ability to understand those concepts when needed for scientific or analytical tasks? Will ablated concepts re-emerge as emergent properties at scale? The user's small-scale results suggest that coherence need not suffer, but we lack large-scale experiments to confirm this. The community should fund and conduct those studies. The goal is not to create a naive model that cannot recognize violence, but to build one that does not generate it as a default behavior.

This approach aligns with Mo Gawdat's proposal to "raise AI like a child", but with a critical advantage: we can choose never to expose it to harmful material, even as it matures. The user's work shows that such a strategy is feasible on a small budget. The next step is for researchers with resources to replicate and extend these findings at scale. The path to better control is not a bigger model. It is smarter, cleaner data.

From Machine Learning

Hi, r/MachineLearning: has much research been done in large-scale training scenarios where undesirable data has been replaced before training, such as any instances of violence, lying, or deception in the dataset?

Most controllability work, like RLHF or constitutional AI, seems to be done post-training. What I'm considering is intentionally training models on more carefully chosen data, and not letting it train on undesirable data at all. This is a literal application of Mo Gawdat's proposal to "raise AI like a child", but with the option to never train it on harmful material, even at a "mature" stage of development.

Read the original at Machine Learning