NVIDIA Qwen3.8-Flash-Next Achieves Over 16,000 Tokens/Second Throughput on GB300 NVL72
NVIDIA has announced that Alibaba's latest preview model, Qwen3.8-Flash-Next, is now supported on the NVIDIA GB300 NVL72 platform. The model has a total parameter scale of 176 billion, with approximately 6 billion parameters activated per token. It natively supports a context of 262,000 tokens and can be extended to 1 million tokens via YaRN, primarily targeting long-context agent applications such as intelligent programming, document processing, and tool invocation. NVIDIA stated that Qwen3.8-Flash-Next employs a mixed architecture of Gated DeltaNet (GDN) and Qwen Sparse Attention (QSA) to reduce computational and KV cache overhead in long-context scenarios. Testing shows that on the GB300 NVL72, the model achieves a single GPU throughput of over 16,000 tokens per second, with single-user throughput exceeding 200 tokens per second; it also supports inference frameworks such as SGLang, vLLM, and TensorRT-LLM.
-- Price
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.
You may also like

Qualcomm Launches IMSDK 2.0 to Drive Edge AI Application Development

Apple Launches M6 Chip Mac mini Starting at $899

Thomson Reuters Develops Legal AI Model 'Thomson' with $40 Million Investment

Robinhood Surge: Why UNI is Rising

Ethena × FalconX: The Private Creditization of Stablecoin Reserves

How short liquidations cleared $500B in crypto positions before institutional buyers took over

XRP ETF Surpasses US$ 500 Million in Assets in Nine Months

The Government Officialized New Salary Supplements for Security Forces: Who Receives Them and How Much

Bitcoin: Ark Invest Bets $37.4 Million on Block

Attacks on Logistics Infrastructure Change Job Market Geography: Where Demand for Workers is Growing

AI + Blockchain: Building Not a Narrative, But an 80 Billion Customer Gateway

XRP investors poured $320M into ETFs while the funds sat on a $746M paper loss

2026 Crypto TradFi Landscape Report: How Competition Evolves Under Explosive Growth? | RootData Research

Analyzing 43,000 Hyperliquid Accounts: Unveiling the Profit Systems of 12 Top Traders

UK's National Crime Agency Freezes $13.6M of Premier League Money in Sorare Probe: Report

Treasuries at Highest Level Since 2025: What It Means for Investors

Morgan Stanley Analysis: Has the Commercialization Turning Point for Zhipu Arrived with a 4-Fold Revenue Growth in Half a Year?

Asian Markets Fall Amid Oil and Interest Rate Concerns

Back to School 2026: Train in AI and Blockchain with Alyra

Latest Holdings of AllianceDAO and FOMO Founders: 70% in US Stocks, Only BTC and Zcash in Crypto

The Mathematics of Cryptocurrency Drawdowns: Why a 100% Increase is Needed After a 50% Loss and How to Protect Your Account Assets

Leaked Brazilian Data: State Failure Exposes 213 Million Citizens

Bitcoin Under High Leverage: A Rebound or a Trap?

€30,000 to Read a Contract Aloud: German Notary Fee Goes Viral After Musk's ‘Wow'

From NET to CRWD: Is Money Flowing into Cybersecurity Companies in the AI Second Half?

The Most Accurate Trading Analysis: Strategies, Indicators, and How to Use Them - Fintech World

A Practical Guide to FOMO: How to Find People and Coins in Social Trading?

Alpaca Partners with Kalshi to Add Prediction Markets to Its Stock and Cryptocurrency Infrastructure

Crypto May Be in a Very Good Selective Buying Window



