25 Aug 2026
Signal Headquarters
Vol. I
No. 237
Reference

What is Qwen?

Qwen

Qwen is a family of large language models developed by Alibaba. The material tracks Qwen’s competitive positioning and technical capabilities, including performance comparisons and integration in speculative decoding.

How it developed

  • Jul 2025 - Brad Gerstner noted that a Qwen 30 billion parameter model released that day was performing as well as GPT-4o.
  • Dec 2025 - Munawar Hayat described an approach called omnidraft that enables transferring draft models for speculative decoding across different model families, demonstrated on Llama 3, Vicuna, and Qwen models.
  • Aug 2026 - Const said that the previous year they produced a model better than Qwen’s best at the 35 billion range, but Qwen launched a superior 35 billion model just before their release.
  • Jul 2026 - Jason Calacanis stated that models from Nvidia’s Neatron family and Google’s Gemma family are not sufficient to replace offerings from Deepseek, Zippu, Moonshot, and Qwen, leading many startups to consider building their own models.

In the evidence

Every line below is attributed to a named speaker.

Contrarian take

Some Qwen models show benchmark improvement when trained with RL on literally random rewards, raising questions about base model quality or evaluation validity.

“We showed that some of the Qwen models actually will improve in if you train them with RL on literally random rewards.”
Nathan Lambert · 10 Dec 2025
By the numbers

A Qwen paper found RL training is approximately 17x more compute-heavy than distillation for small models.

“The Quen paper showed that the RL was like kind of 17 times more compute.”
Ross Taylor · 29 Jul 2025
Company & tool watch

OmniDraft transfers speculative decoding draft models across different model families (Llama 3, Vicuna, Qwen) by using an n-gram cache to bridge tokenizer differences, enabling cross-family inference acceleration.

“It's called omnidraft so basically it enables you to transfer the draft model train for speculative decoding across different families of networks and they I guess they show it on Llama 3, Vikuna and Quen models and they come up with some with an with an interesting approach and Graham cache the difference the main the main difference is across families of models we have different tokenizers and how do you make sure that one tokenizer corresponds to another one and engram cache basically bridges that gap.”
Munawar Hayat · 9 Dec 2025
Contrarian take

A Bittensor subnet built a model that outperformed Qwen's best 35B model, but Qwen shipped a superior 35B update just before the Bittensor launch, illustrating the speed disadvantage decentralized teams face against centralized incumbents.

“Last year we did produce a model that was better than Quen's best model at the 35 billion range. And just before we launched it, they launched another 35 billion model that was better than ours.”
Const · 17 Aug 2026
Company & tool watch

Qwen maintains a daily-updated preview API endpoint, enabling near-continuous iteration that gives it outsized influence over the research ecosystem.

“They have an endpoint which you can use for a preview version and they've updated this endpoint daily.”
Nathan Lambert · 22 Jul 2026
By the numbers

Open reasoning models including Qwen exhibit repetition meltdowns in roughly 0.01% or more of samples, where paragraphs repeat 50 to 100 times before the model recovers.

“0.01% or even more. So, like one in a thousand or one in ten thousand samples from Qwen have an mental break where the model will just repeat paragraphs like 50 or 100 times.”
Nathan Lambert · 10 Dec 2025
By the numbers

Alibaba's Qwen 30 billion parameter model is reported to perform on par with GPT-4o.

“There was a release today of a quen you know 30 billion parameter model which is performing as good as GPT40.”
Brad Gerstner · 31 Jul 2025
Contrarian take

US open-source models from Nvidia (Nemotron) and Google (Gemma) are not sufficient substitutes for Chinese open-weight models like Deepseek, Qwen, and Moonshot, leaving startups avoiding paid APIs with no reliable fallback.

“The models from Nvidia, the Neatron family, that the models from Google's Gemma family aren't sufficient to replace what we have from Deepseek, from Zippu, from Moonshot, from Quinn, and we're going to end up in a place where a lot of people are betting on rolling their own, especially startups that don't want to pay open margin for them, and they're just not going to have something to fall back on.”
Jason Calacanis · 9 Jul 2026
Signal Headquarters · reference note, compiled from attributed expert discussion. Last updated 2026-08-18.