The reporting

1 articles

Topic: research

Chart showing where constitutional midtraining alignment gains remained after later training

NEWS AI Science

Oxford Researchers Put AI Principles Earlier in Training, and Some of Them Stuck

A 120-billion-parameter experiment found that adding constitutional content before conventional safety tuning produced alignment gains that survived later training. The strongest results are promising, but narrower than a permanent solution to AI alignment.