Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Architecture thread! Afaict they continue to use gated attention + delta net, which was also adopted+adapted by K3, but im surprised theres no improvements to the residual stream (deepseek are using manifold hyper-connections, kimi have attention residuals) ?

Perf improvements seem to all come from training?



As was the case with GLM 5.3, it seems that there is still much juice to be squeezed from post-training




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: