AlibabaがQwen4の設計を先行公開 — バックボーン125B・合計180BのマルチモーダルMoEモデル、推論時の活性化は6Bのみで訓練コストは前世代の1/9
DRANK

8月27日、MarkTechPostが「Alibaba's Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture」と題した記事を公開した。AlibabaのQwenチームがQwen4アーキテクチャの先行実装として位置づけるオープンウェイトのマルチモーダルMoEモデル「Qwen3.8-Flash-Next」の詳細が明らかになった。

by @tf_official
Related Topics: Apache HTTP Server Machine Learning Deep Learning