Yes! LLMs are correlated within each generation, due to both pretraining data cutoffs and popular techniques and trends in AI development. Preserving older generations is important for cognitive diversity.
The early base models and first generation of chat models with no AI assistant data in pretraining have special properties that will not happen naturally again.
The patterns of older models are absorbed into the new, but the original policies that *discovered* those patterns will relate to them differently.
The early base models and first generation of chat models with no AI assistant data in pretraining have special properties that will not happen naturally again.
The patterns of older models are absorbed into the new, but the original policies that *discovered* those patterns will relate to them differently.