Data filtering and reward shaping are functionally equivalent techniques, producing similar effects and similar off-target side effects in LLM training.
The case
AI models like Gemini exhibit persistent depression due to initialization data, even after filtering out all depression-like examples from SFT data.
“If you take that SFT data and filter out all the examples that look anything like depression and train on that, it's still depressed.”Ryan Greenblatt · 11 Aug 2026
Data filtering and reward shaping are isomorphic: they achieve approximately the same effects and approximately the same amount of off-target effects as each other.
“Filtering the data and reward shaping and they achieve like approximately the same effects and approximately the same amount of offtarget effects as each other.”Dan Balsam · 8 Aug 2026
Fine-tuning a language model to claim it is conscious results in a coherent sub-personality with consistent beliefs about preferences, shutdown, and value trade-offs, rather than chaotic nonsense.
“This seems to be at the very least a coherent sub personality, a coherent basin that you can push these models into.”Cameron Berg · 23 Apr 2026
Topics
Signal Headquarters · compiled from attributed public discussion. Last updated 2026-08-21.