Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The idea that they are entirely trained on human intelligence is already outdated. Yes, the earlier models relied heavily on RLHF and human curated data but we have since moved on to synthetic data produced by the models themselves and reinforcement learning with verifiable rewards (RLVR).
 help



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: