Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Yes there have been several research projects like these that have fully open sourced their training data (mostly coming from Europe). But that's all they amount to. Experiments and projects.

The labs that are actually competing with frontier models are using data would usually be a violation of copyright to release openly.

> Chinese models these days don't even release pre-trained weights anymore; all you get now is the finished post-trained product.

No? That's absolutely not true. Qwen, GLM, Kimi, DeepSeek, etc all consistently release both the post-trained "Instruct/Chat" versions and the underlying "Base" (pre-trained) weights.

Which specific Chinese models are you thinking about?



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: