Skip to content

test: add Qwen3 Megatron baseline - #15

Open
JYMiracle305 wants to merge 2 commits into
feat/add-llama3-megatron-baselinefrom
feat/add-qwen3-megatron-baseline
Open

JYMiracle305 wants to merge 2 commits into
feat/add-llama3-megatron-baselinefrom
feat/add-qwen3-megatron-baseline

Conversation

@JYMiracle305

@JYMiracle305 JYMiracle305 commented Aug 21, 2026 •

Copy link
Copy Markdown
Collaborator

概述

本 PR 在 Megatron shadow baseline 中新增 Qwen3-8B,用于让 InfiniTrain 和 Megatron-LM 从同一份 LLMC v4 初始权重、同一串 token 和同一组训练参数启动,并逐步比较 loss。

该 PR 基于 #14,复用 baseline/megatron/common/ 中的数据转换和比较工具。

目录结构

baseline/megatron/models/qwen3/
├── README.md
├── prepare/
│   ├── convert_hf_qwen3_to_llmc.py  # Hugging Face 权重转共享 LLMC v4,用于infiniTrain训练读取
│   └── prepare_dataset.sh            # token 转换成 Megatron indexed dataset
├── adapter/
│   └── llmc_loader.py                # LLMC 读取、TP 切分和 Megatron 参数映射
├── train/
│   ├── pretrain.py                     # 接入 Megatron upstream 训练流程
│   └── run_training.sh               # Qwen3 配置与训练入口
└── compare/
    └── compare_loss.sh               # 比较 InfiniTrain/Megatron 逐步 loss

运行时产生的数据集、cache、日志和比较结果放在 baseline/megatron/artifacts/qwen3/,不进入 Git 提交。

数据流

Hugging Face Qwen3-8B
        │
        ▼
convert_hf_qwen3_to_llmc.py
        │
        ▼
完整 LLMC v4 权重
├── InfiniTrain:C++ loader 直接读取
└── Megatron:llmc_loader.py 读取并按 TP 映射

同一份 token stream
├── InfiniTrain:直接读取
└── Megatron:转换为 indexed dataset 后读取

Qwen3 adapter 会校验 LLMC v4 header 和文件大小,处理 Q/K Norm、GQA QKV 布局、SwiGLU gate/up 顺序以及 TP 权重切分。

验证方式

QWEN3_INPUT_BIN=/path/to/llmc_tokens.bin \
  bash baseline/megatron/models/qwen3/prepare/prepare_dataset.sh

QWEN3_LLMC_FILEPATH=/path/to/qwen3-8b-fp32.llmc \
DTYPE=float32 NPROC_PER_NODE=8 TP=8 PP=1 \
  bash baseline/megatron/models/qwen3/train/run_training.sh

INFINITRAIN_LOG=/path/to/result_infinitrain_training.log \
  bash baseline/megatron/models/qwen3/compare/compare_loss.sh

当前 baseline 面向 Dense Qwen/Qwen3-8B。

@JYMiracle305 JYMiracle305 changed the title test: add Qwen3 Megatron baseline [WIP] test: add Qwen3 Megatron baseline Aug 24, 2026
@JYMiracle305
JYMiracle305 force-pushed the feat/add-qwen3-megatron-baseline branch from 6acc435 to d4f6951 Compare September 4, 2026 09:22
@JYMiracle305 JYMiracle305 changed the title [WIP] test: add Qwen3 Megatron baseline test: add Qwen3 Megatron baseline Sep 9, 2026
@JYMiracle305
JYMiracle305 force-pushed the feat/add-qwen3-megatron-baseline branch from d4f6951 to ffb7e14 Compare September 9, 2026 07:07
@JYMiracle305
JYMiracle305 force-pushed the feat/add-qwen3-megatron-baseline branch from e1989e4 to 95decd9 Compare September 24, 2026 01:52
@JYMiracle305
JYMiracle305 force-pushed the feat/add-qwen3-megatron-baseline branch 2 times, most recently from fb6349d to 200e9dd Compare September 24, 2026 05:50
parser.add_argument("--include-cases", help="comma-separated log basenames to compare")
parser.add_argument("--prefer-autocast-bfloat16", action="store_true", help="prefer historical Megatron autocast BF16 log names")
parser.add_argument("--threshold-fp32", type=float, default=1e-5)
parser.add_argument("--threshold-sp-fp32", type=float, default=2e-2)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这里可以验证下,经过 InfiniTensor/InfiniTrain#224 的修复后,Qwen3 的 sp 测例阈值是否仍然需要特判。

@JYMiracle305
JYMiracle305 force-pushed the feat/add-qwen3-megatron-baseline branch from 200e9dd to 46e6c0b Compare October 10, 2026 03:25
@JYMiracle305
JYMiracle305 force-pushed the feat/add-qwen3-megatron-baseline branch 3 times, most recently from fb5c1b7 to b0c945b Compare October 10, 2026 10:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants