Repository navigation
test: add GPT2 Megatron baseline - #17
Open
JYMiracle305 wants to merge 3 commits into
Open
JYMiracle305 wants to merge 3 commits into
JYMiracle305 wants to merge 3 commits into
Conversation
JYMiracle305
force-pushed
the
feat/add-gpt2-megatron-baseline
branch
from
September 9, 2026 07:07
f5073fe to
178587c
Compare
JYMiracle305
added this pull request to stack #16
September 9, 2026 07:09
Collaborator
Author
JYMiracle305
force-pushed
the
feat/add-gpt2-megatron-baseline
branch
from
September 24, 2026 01:52
b48ea88 to
544663c
Compare
Collaborator
|
当前截图中的 failed loss cases 可能与 InfiniTensor/InfiniTrain#224 修复的问题有关,建议基于 InfiniTrain 最新 master 重新对比后,更新精度对比情况截图。 |
JYMiracle305
force-pushed
the
feat/add-gpt2-megatron-baseline
branch
2 times, most recently
from
October 10, 2026 06:04
0187a8b to
19f337e
Compare
JYMiracle305
force-pushed
the
feat/add-gpt2-megatron-baseline
branch
from
October 10, 2026 10:02
19f337e to
92220e2
Compare
JYMiracle305
force-pushed
the
feat/add-gpt2-megatron-baseline
branch
from
October 10, 2026 10:07
92220e2 to
587b398
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

概述
本 PR 在现有 Megatron shadow baseline 中接入 GPT-2。InfiniTrain、PyTorch 和 Megatron-LM 使用相同 LLMC FP32 初始权重、相同 Tiny Shakespeare token 顺序和相同训练参数,用于比较逐步 loss 与标准训练过程中的吞吐。
GPT-2 baseline 沿用 Llama3 的目录组织、BF16 autocast 语义、10 步基础用例和批量日志比较方式;不修改
third_party/Megatron-LM。运行生成的 dataset、cache、日志和比较结果位于baseline/megatron/artifacts/,不会提交到仓库。Stack 关系
本 PR 的 base 是
feat/add-qwen3-megatron-baseline。目录结构
根目录
README.md同步增加 GPT-2 Megatron 快速入口。报告、日志、cache、artifacts 和__pycache__均未包含在提交中。数据与模型对齐
GPT-2 参数适配会校验 LLMC 文件头、恢复运行时词表大小,并将 QKV 按 attention head 交错为 Megatron MCore 所需布局。当前 LLMC 直接加载支持 TP=1、PP=1;数据并行用例支持单卡和 8 卡。
基础用例
gpt2_1gpt2_1_bfloat16gpt2_2gpt2_2_bfloat16gpt2_3gpt2_3_bfloat16BF16 基础用例使用
DTYPE=autocast_bfloat16:参数保持 FP32,forward 位于 PyTorch BF16 autocast 上下文中,与 InfiniTrain/PyTorch 的精度策略对齐。Megatron 原生DTYPE=bfloat16仍可用于额外诊断,但不属于默认基础矩阵。FP32 默认关闭 TF32。精度与性能统计
2e-31e-2train_step内tok/s和InfiniTrain/Megatron百分比验证结果
使用 InfiniTrain master 日志
scripts/20260904/master_e403ed4/logs/basic对比:gpt2_11.15e-32e-3gpt2_1_bfloat162.11e-21e-2gpt2_27.03e-42e-3gpt2_2_bfloat169.23e-31e-2gpt2_36.99e-42e-3gpt2_3_bfloat169.23e-31e-2六个 Megatron 训练用例均运行成功;loss 为
5/6 PASS,TPS 为6/6 cases compared,无日志缺失或解析错误。gpt2_1_bfloat16是 256-token 小 batch,用例对 BF16 舍入和运行版本较敏感。其 Megatron/InfiniTrain 最大差异为2.11e-2,此前对应 PyTorch/InfiniTrain 对比也达到约2.30e-2,属于同一数量级。因此保留1e-2标准阈值并明确显示 FAIL,不通过放宽阈值隐藏差异。验证方式
准备数据并运行全部六个 Megatron 用例:
只运行选定用例:
比较单个用例并输出逐步 JSON:
直接比较 InfiniTrain 标准运行目录和 Megatron 日志目录:
loss 脚本明确列出 PASS、FAIL、缺失日志和总计;TPS 脚本仅输出平均吞吐、比例、解析错误和缺失日志。任一解析错误或缺失日志都会返回非零退出码。
检查
bash -ngit diff --check通过