Repository navigation
test: add Llama3 Megatron baseline - #14
Open
JYMiracle305 wants to merge 7 commits into
Open
JYMiracle305 wants to merge 7 commits into
JYMiracle305 wants to merge 7 commits into
Conversation
JYMiracle305
commented
Aug 31, 2026
|
|
||
|
|
||
| if __name__ == "__main__": | ||
| gpt_dataset._build_shuffle_index = ordered_shuffle_index |
Collaborator
Author
There was a problem hiding this comment.
这里将 Megatron 默认的 shuffle 索引替换为顺序索引
JYMiracle305
force-pushed
the
feat/add-llama3-megatron-baseline
branch
from
September 4, 2026 09:08
050b4e5 to
47a38b9
Compare
Collaborator
Author
JYMiracle305
force-pushed
the
feat/add-llama3-megatron-baseline
branch
from
September 4, 2026 09:15
47a38b9 to
3a5990b
Compare
JYMiracle305
force-pushed
the
feat/add-llama3-megatron-baseline
branch
3 times, most recently
from
September 24, 2026 02:03
f4e0804 to
e8aaa52
Compare
kilinchange
reviewed
Oct 9, 2026
|
|
||
| Generated files under `baseline/megatron/artifacts/` are not committed. | ||
|
|
||
| ## Llama3 Megatron baseline |
Collaborator
There was a problem hiding this comment.
建议将 Llama3 Megatron baseline 的具体使用说明放在对应目录的 README 中,仓库根目录的 README 主要保留项目整体介绍及各类 baseline 的使用入口。
目前根目录 README 只详细介绍了 Llama3 Megatron baseline,未涉及其他模型及 PyTorch baseline,文档组织上不太统一,也不利于后续扩展。
如果希望统一在根目录维护使用说明,也建议一并补充其他模型及 PyTorch baseline 的相关内容。
Collaborator
Author
There was a problem hiding this comment.
这里先删掉,各个模型目录下现在已有README
JYMiracle305
force-pushed
the
feat/add-llama3-megatron-baseline
branch
from
October 10, 2026 03:25
e8aaa52 to
4ee7794
Compare
JYMiracle305
force-pushed
the
feat/add-llama3-megatron-baseline
branch
from
October 10, 2026 06:04
4ee7794 to
3725c37
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.


概述
本 PR 在
InfiniTrain-Test中建立按模型组织的 Megatron shadow baseline,并先接入 Llama3。InfiniTrain、PyTorch 和 Megatron-LM 使用相同 LLMC 初始权重、相同 token 顺序和相同训练参数,以比较逐步 loss 与训练吞吐。现有
baseline/pytorch用例保持不变;Megatron baseline 作为独立目录并行验证,不替换原有 PyTorch baseline,也不修改third_party/Megatron-LM。目录结构
运行生成的数据集、cache、日志和比较结果放在
baseline/megatron/artifacts/,由.gitignore排除,不进入仓库。根目录README.md已补充仓库结构和 Llama3 Megatron 快速入口。数据与模型对齐
Llama3 adapter 校验 LLMC header 和完整文件大小,将 block Q/K/V 重排为 Megatron GQA 布局,并按 Megatron SwiGLU 顺序打包 gate/up 权重。
基础用例
llama3_1llama3_1_bfloat16llama3_2llama3_2_bfloat16llama3_3llama3_3_bfloat16BF16 基础用例保留 FP32 参数,在 forward 中使用 PyTorch BF16 autocast,与 InfiniTrain/PyTorch 的精度策略对齐。Megatron 原生
DTYPE=bfloat16仍可用于额外诊断,但不属于默认基础矩阵。精度与性能统计
1e-51e-2train_step内与配套 InfiniTrain 日志比较时,六个基础用例 loss 均通过各自阈值。BF16 autocast 相比 Megatron 原生 BF16 明显改善数值一致性;256-token BF16 小 batch 对不同运行版本更敏感,其他 InfiniTrain 运行记录中可能出现略高于
1e-2的单步差异。验证方式
准备数据并运行全部六个 Megatron 用例:
只运行选定用例:
比较单个用例并输出逐步 JSON:
直接比较 InfiniTrain 标准运行目录和 Megatron 日志目录:
loss 脚本明确列出 PASS、FAIL、缺失日志和总计;TPS 脚本仅输出平均吞吐、比例、解析错误和缺失日志。任一解析错误或缺失日志都会返回非零退出码。
后续扩展
common/只放跨模型复用的数据转换和单例诊断工具,scripts/放跨模型批量比较工具;新增模型放在models/<model>/下,并分别维护prepare、adapter、train和compare。