Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I did a quick search for "llama" and didn't find anywhere they outright state they just fine-tuned some llama weights.

Is it possible that they based their model architecture on the llama model architecture? Rather than just fine-tuned already training llama weights? In that case, they'd still have to do "bottoms up" training.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: