Super crazy that their GPQA scores are that high considering they tested at 0-shot. I almost worry there might be some leakage.
Super excited for what the big Llama-3 is going to bring to the table.
Flaky91 on
Some interesting notes.
* 8b parameter version and 70b parameter version.
* decoder only architecture.
* Text in to text out only on the models (currently).
* Plans to release multimodal versions of llama 3 later
* Plans to release larger context windows later.
* It generally sounds like they’re going for an iterative release.
* Pretrained on 15 trillion tokens.
* Trained on 2 24k GPU clusters.
* New more efficient tokenizer and a vocabulary of 128k tokens.
* Have versions still in training internally at over 400b parameters.
* Created an internal evaluation that was never given to the modeling team in order to avoid overfitting.
oddmetre on
How does it compare with GPT-4 in the “Instruct Human” evaluation? They only compared with GPT 3.5 according to the diagram but maybe I missed it in the article
Leave A Reply
Du musst angemeldet sein, um einen Kommentar abzugeben.
3 Kommentare
Really impressive results out of Meta here.
Super crazy that their GPQA scores are that high considering they tested at 0-shot. I almost worry there might be some leakage.
Super excited for what the big Llama-3 is going to bring to the table.
Some interesting notes.
* 8b parameter version and 70b parameter version.
* decoder only architecture.
* Text in to text out only on the models (currently).
* Plans to release multimodal versions of llama 3 later
* Plans to release larger context windows later.
* It generally sounds like they’re going for an iterative release.
* Pretrained on 15 trillion tokens.
* Trained on 2 24k GPU clusters.
* New more efficient tokenizer and a vocabulary of 128k tokens.
* Have versions still in training internally at over 400b parameters.
* Created an internal evaluation that was never given to the modeling team in order to avoid overfitting.
How does it compare with GPT-4 in the “Instruct Human” evaluation? They only compared with GPT 3.5 according to the diagram but maybe I missed it in the article