Skip to content

feat(models): add Qwen3.5-122B-A10B-FP8 candidate - #307

Draft
krisztian-gajdar wants to merge 1 commit into
mainfrom
codex/onboard-qwen35-122b
Draft

krisztian-gajdar wants to merge 1 commit into
mainfrom
codex/onboard-qwen35-122b

Conversation

@krisztian-gajdar

Copy link
Copy Markdown
Contributor

Adds Qwen/Qwen3.5-122B-A10B-FP8 to the existing SGLang adapter, pinned to Hugging Face revision a099dee70ccfcd8d5dda56aaa0b60cb8ecadabc9.

The candidate uses native FP8 weights, BF16 compute, tensor parallelism across two GPUs, and a bounded 32K non-thinking context with a 4096-token output cap. It includes an explicit rtx-pro-6000x2 profile and a matching non-speculative grammar profile. Tool parsing and sampling defaults follow the pinned model card.

Validation: canonical model configuration and adapter-option checks pass; catalog and tensor-parallel tests report 294 passed and 33 skipped. GPU acceptance remains pending, so this PR stays draft.

@coderabbitai

coderabbitai Bot commented Sep 18, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Comment @coderabbitai help to get the list of available commands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant