A helpful paper with the full recipe Cerebras uses to train LLMs and their proce...

		jwan584 on Sept 22, 2023 \| parent \| context \| favorite \| on: BTLM-3B-8K: 7B Parameter Performance in a 3B Param... A helpful paper with the full recipe Cerebras uses to train LLMs and their process including: - Extensively deduplicated dataset (SlimPajama) - Hyperparameter search using muP - Variable sequence length training + ALiBi - Aggressive LR decay