GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop
-
This post did not contain any content.
-
This post did not contain any content.
[citation needed]
-
[citation needed]
This is for qwen 3.5, not 3.6, but the vram requirements are a little out there for most:
Qwen 3.5 27B VRAM Requirements — Dense Model Hardware Guide (Q4/Q5/Q6/Q8) | Will It Run AI Blog
Qwen 3.5 27B needs ~16.5 GB at Q4_K_M on RTX 4090. See also the newer Qwen3.6-27B (April 22, 2026) which needs 16.8 GB Q4 and beats it on coding benchmarks.
Will It Run AI (willitrunai.com)
16 to 55GB...
-
This is for qwen 3.5, not 3.6, but the vram requirements are a little out there for most:
Qwen 3.5 27B VRAM Requirements — Dense Model Hardware Guide (Q4/Q5/Q6/Q8) | Will It Run AI Blog
Qwen 3.5 27B needs ~16.5 GB at Q4_K_M on RTX 4090. See also the newer Qwen3.6-27B (April 22, 2026) which needs 16.8 GB Q4 and beats it on coding benchmarks.
Will It Run AI (willitrunai.com)
16 to 55GB...
I mean that's high but not unreasonable, especially with unified ram machines
-
I mean that's high but not unreasonable, especially with unified ram machines
and it's even lower with MTP https://huggingface.co/unsloth/Qwen3.6-27B-MTP-GGUF
-
I mean that's high but not unreasonable, especially with unified ram machines
Little outside my budget personally, but I've definitely been itching to pull the trigger on a nVidia P40
-
This post did not contain any content.
Which one of them made these overlapping bar labels?
-
This post did not contain any content.
It’s a bit misleading.
Qwen 27B has way less “world knowledge” than GPT-5. Ask it random trivia without internet search access, and GPT would know waaay more.
This is generally true of small vs large models.
…But honestly, Qwen 27B is better at tool use or agentic stuff. It’s hyper optimized for just that and coding assistance, basically.
This is often true of old vs new. Most newer models have hyper focused on agents/coding, often to the detriment of other use cases.
Quantization for practically running Qwen 27V also has an impact. A off-the-shelf Q4_K_M is not the same as the unquantized weights in real-world use, or even an “optimized” quantization like a custom exl3.
-
It’s a bit misleading.
Qwen 27B has way less “world knowledge” than GPT-5. Ask it random trivia without internet search access, and GPT would know waaay more.
This is generally true of small vs large models.
…But honestly, Qwen 27B is better at tool use or agentic stuff. It’s hyper optimized for just that and coding assistance, basically.
This is often true of old vs new. Most newer models have hyper focused on agents/coding, often to the detriment of other use cases.
Quantization for practically running Qwen 27V also has an impact. A off-the-shelf Q4_K_M is not the same as the unquantized weights in real-world use, or even an “optimized” quantization like a custom exl3.
I mean baking knowledge into a model isn't really all that useful to begin with. Just download wikipedia locally and have it access it through tool use, it's way more efficient and more accurate. And yeah, I find Q6 tends to be the sweet spot where it's close enough to full 16 bit in performance, but doesn't chew up too much memory.
-
This is for qwen 3.5, not 3.6, but the vram requirements are a little out there for most:
Qwen 3.5 27B VRAM Requirements — Dense Model Hardware Guide (Q4/Q5/Q6/Q8) | Will It Run AI Blog
Qwen 3.5 27B needs ~16.5 GB at Q4_K_M on RTX 4090. See also the newer Qwen3.6-27B (April 22, 2026) which needs 16.8 GB Q4 and beats it on coding benchmarks.
Will It Run AI (willitrunai.com)
16 to 55GB...
I can run both the 27B and 35B version of Qwen3.6 on my 10GB VRAM 3080, and they run ok. You don't need to put the entire model in VRAM, even if it's probably beneficial.
-
It’s a bit misleading.
Qwen 27B has way less “world knowledge” than GPT-5. Ask it random trivia without internet search access, and GPT would know waaay more.
This is generally true of small vs large models.
…But honestly, Qwen 27B is better at tool use or agentic stuff. It’s hyper optimized for just that and coding assistance, basically.
This is often true of old vs new. Most newer models have hyper focused on agents/coding, often to the detriment of other use cases.
Quantization for practically running Qwen 27V also has an impact. A off-the-shelf Q4_K_M is not the same as the unquantized weights in real-world use, or even an “optimized” quantization like a custom exl3.
What version do you use and how do you run Qwen3.6? I've played around a bit with the Q4 version in LM studio +Zed, but I was not happy with the results. It looses track very often and often enters infinite loops or just stops...
-
It’s a bit misleading.
Qwen 27B has way less “world knowledge” than GPT-5. Ask it random trivia without internet search access, and GPT would know waaay more.
This is generally true of small vs large models.
…But honestly, Qwen 27B is better at tool use or agentic stuff. It’s hyper optimized for just that and coding assistance, basically.
This is often true of old vs new. Most newer models have hyper focused on agents/coding, often to the detriment of other use cases.
Quantization for practically running Qwen 27V also has an impact. A off-the-shelf Q4_K_M is not the same as the unquantized weights in real-world use, or even an “optimized” quantization like a custom exl3.
I sort of get it because of you know you're only using the model when you have Internet and want it to search then just let it.
-
Which one of them made these overlapping bar labels?
Hey, a lot of people sucked making matplotlib plots before LLMs existed.
-
This post did not contain any content.
So which one tried to cheat its way into the test?
Llama is one for sure
-
I can run both the 27B and 35B version of Qwen3.6 on my 10GB VRAM 3080, and they run ok. You don't need to put the entire model in VRAM, even if it's probably beneficial.
Full quantisation? I've only got a 8GB 3070, but I'll give it a go
Edit: Tried the unsloth/qwen3.6 with llama.CPP, and it failed to allocate a 26GB Vulcan buffer and died. Dunno what magic your using, no luck for me though

-
This post did not contain any content.
what kinda hardware do you need to "run this on your desktop"?
-
what kinda hardware do you need to "run this on your desktop"?
You can run it on systems with as little as 32 GBs of RAM.
-
This post did not contain any content.
Those models are for general use. If you have business use case and data related to it you can finetune model for specific use that will outperform all of frontier models and run at fraction of cost.
-
This post did not contain any content.
Kimi is better. Waiting for it to appear on Ollama Cloud.
-
You can run it on systems with as little as 32 GBs of RAM.
You don't want to run these models in RAM. I started using them on an RTX 3060-12Gb VRAM, and quickly added another, for 24Gb RAM. I also have 64 Gb RAM. Now it works well. Not lightning fast, but quite useable. VRAM is the key.
Ciao! Sembra che tu sia interessato a questa conversazione, ma non hai ancora un account.
Stanco di dover scorrere gli stessi post a ogni visita? Quando registri un account, tornerai sempre esattamente dove eri rimasto e potrai scegliere di essere avvisato delle nuove risposte (tramite email o notifica push). Potrai anche salvare segnalibri e votare i post per mostrare il tuo apprezzamento agli altri membri della comunità.
Con il tuo contributo, questo post potrebbe essere ancora migliore 💗
Registrati Accedi
Citiverse è un progetto che si basa su NodeBB ed è federato! | Categorie federate | Chat | 📱 Installa web app o APK | 🧡 Donazioni | Privacy Policy