Huawei opensouced openPangu-2.0-Pro with 505B total parameters and 18B activated parameters; context length 512k, 34T tokens pretraining, trained via unified SFT with slow/fast thinking, multiple RL specialists.
Read the original at old.reddit.com→openPangu-2.0-Pro is an MoE model trained on Ascend. The model has 505B total parameters and 18B activated parameters. Its context length is 512k. The total pretraining data contains 34T tokens. During Post-training,...
Original headline: "Huawei opensouced openPangu-2.0-Pro, 505B-A18B"
Coverage timeline
- Jul 31, 06:47 UTC r/LocalLLaMA lead source Huawei opensouced openPangu-2.0-Pro, 505B-A18B