<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop]]></title><description><![CDATA[<em>This post did not contain any content.</em>]]></description><link>https://citiverse.it/topic/558ba0bb-f4b6-496c-b9e1-6256a1fd2a13/gpt-5-the-world-best-model-just-1-year-ago-is-today-inferior-to-qwen3.6-27b-that-you-can-run-on-your-desktop</link><generator>RSS for Node</generator><lastBuildDate>Tue, 18 Aug 2026 10:30:43 GMT</lastBuildDate><atom:link href="https://citiverse.it/topic/558ba0bb-f4b6-496c-b9e1-6256a1fd2a13.rss" rel="self" type="application/rss+xml"/><pubDate>Tue, 28 Jul 2026 22:35:08 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Sat, 08 Aug 2026 22:15:09 GMT]]></title><description><![CDATA[<p dir="auto">Facts! Thank you!</p>
]]></description><link>https://citiverse.it/post/https://programming.dev/comment/25341017</link><guid isPermaLink="true">https://citiverse.it/post/https://programming.dev/comment/25341017</guid><dc:creator><![CDATA[racketman@programming.dev]]></dc:creator><pubDate>Sat, 08 Aug 2026 22:15:09 GMT</pubDate></item><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Thu, 30 Jul 2026 18:31:11 GMT]]></title><description><![CDATA[<p dir="auto"><img src="https://citiverse.it/assets/plugins/nodebb-plugin-emoji/emoji/android/1f644.png?v=b09b169718a" class="not-responsive emoji emoji-android emoji--face_with_rolling_eyes" style="height:23px;width:auto;vertical-align:middle" title="🙄" alt="🙄" /></p>
<p dir="auto">Is it still telling people to put glue on their pizza and convincing teenagers to kill themselves?</p>
]]></description><link>https://citiverse.it/post/https://lemmy.ml/comment/26984572</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.ml/comment/26984572</guid><dc:creator><![CDATA[nonconfrontational@lemmy.ml]]></dc:creator><pubDate>Thu, 30 Jul 2026 18:31:11 GMT</pubDate></item><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Thu, 30 Jul 2026 18:10:11 GMT]]></title><description><![CDATA[<p dir="auto">An important caveat to that explanation is that the mixture of experts gets re-evaluated for every single token in the input sequence and not for high level tasks as in the example. There will be certain "experts" for looking at indentation tokens, or ones that look at specific word beginnings/prefixes etc.</p>
]]></description><link>https://citiverse.it/post/https://discuss.tchncs.de/comment/27266229</link><guid isPermaLink="true">https://citiverse.it/post/https://discuss.tchncs.de/comment/27266229</guid><dc:creator><![CDATA[starfighter@discuss.tchncs.de]]></dc:creator><pubDate>Thu, 30 Jul 2026 18:10:11 GMT</pubDate></item><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Thu, 30 Jul 2026 16:18:49 GMT]]></title><description><![CDATA[<p dir="auto">Honestly, I think the most reasonable approach is just to see what other people's experience is like and which models are well regarded, then try them out and see which one is the best fit for what you're doing. You might not even need the top performing one necessarily, and speed or lower resource usage might be a bigger factor.</p>
]]></description><link>https://citiverse.it/post/https://lemmy.ml/comment/26982316</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.ml/comment/26982316</guid><dc:creator><![CDATA[yogthos@lemmy.ml]]></dc:creator><pubDate>Thu, 30 Jul 2026 16:18:49 GMT</pubDate></item><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Thu, 30 Jul 2026 16:03:25 GMT]]></title><description><![CDATA[<p dir="auto">i see what your saying. i didnt mean to discredit standard benchmarks entirely.<br />
i guess its obvious that it measures capability regardless of imprecision.<br />
2 major proposed changes:<br />
**first, i dont really know. aside from saying "benchmark your own prompt+usecase"<br />
a proposed plan:</p>
<ul>
<li>approach one: pay attention and credit new or improved architecture designs and research.</li>
<li>approach two: spend more attention on benchmarks. especially specific benchmarks ( that are not focused  with industrial domain tasks.) **domain task pursuit, is useful!.. but it depends on if your interest align to popular domains.</li>
<li>approach three: if willing to utilize remotely hosted models. rating should also take in consideration.. tools and everything else: websearch performance, RAG performance, smooth interface, pref/balance between speed vs comprehensiveness, cost (if relevant), etc.. .</li>
</ul>
]]></description><link>https://citiverse.it/post/https://lemmy.ml/comment/26982079</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.ml/comment/26982079</guid><dc:creator><![CDATA[leanleft@lemmy.ml]]></dc:creator><pubDate>Thu, 30 Jul 2026 16:03:25 GMT</pubDate></item><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Thu, 30 Jul 2026 11:13:35 GMT]]></title><description><![CDATA[<p dir="auto">That's why you couple it with your own, self-hosted yacy or searxng instance. Embedded world knowledge does not help a model if it becomes outdated. I just let my agent research, embed that knowledge to a little Qdrant server, so other servers are not bothered again and pull the information from there when needed again. With a little RAG you can have GPT at home.</p>
]]></description><link>https://citiverse.it/post/https://lemmy.zip/comment/27937850</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.zip/comment/27937850</guid><dc:creator><![CDATA[userentity@lemmy.zip]]></dc:creator><pubDate>Thu, 30 Jul 2026 11:13:35 GMT</pubDate></item><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Thu, 30 Jul 2026 11:08:38 GMT]]></title><description><![CDATA[<p dir="auto">I had to get mine from ebay and wait a couple weeks as they came from China too. Also I recommend having a 3D-Printer and some Blower-Fans on hand as you will either have to buy or print your own fan shroud for these server cards.</p>
]]></description><link>https://citiverse.it/post/https://lemmy.zip/comment/27937786</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.zip/comment/27937786</guid><dc:creator><![CDATA[userentity@lemmy.zip]]></dc:creator><pubDate>Thu, 30 Jul 2026 11:08:38 GMT</pubDate></item><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Thu, 30 Jul 2026 10:45:46 GMT]]></title><description><![CDATA[<p dir="auto">thanks, good stuff! sadly, got none of these locally available.</p>
]]></description><link>https://citiverse.it/post/https://programming.dev/comment/25172995</link><guid isPermaLink="true">https://citiverse.it/post/https://programming.dev/comment/25172995</guid><dc:creator><![CDATA[yuman@programming.dev]]></dc:creator><pubDate>Thu, 30 Jul 2026 10:45:46 GMT</pubDate></item><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Thu, 30 Jul 2026 11:09:57 GMT]]></title><description><![CDATA[<p dir="auto">A P40 with 24GB is ~150€, a V100 (32GB) is ~600€. Both of these fit Qwen3.6 27B (The P40 is about 3x slower though). The V100 even fits 400k context with a Q4 KV-Cache , which means you can have two slots for parallel processing (llama-cpp). You don't even have to use system memory.<br />
One of my inference servers is running with 8gb of DDR3 and a 2nd Gen i7, so my old hardware has a good use again.</p>
]]></description><link>https://citiverse.it/post/https://lemmy.zip/comment/27937424</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.zip/comment/27937424</guid><dc:creator><![CDATA[userentity@lemmy.zip]]></dc:creator><pubDate>Thu, 30 Jul 2026 11:09:57 GMT</pubDate></item><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Thu, 30 Jul 2026 02:46:49 GMT]]></title><description><![CDATA[<p dir="auto">Running model that is good at everything require huge amount of energy and huge data center. Those models are mixture of experts. Latest Kimi K3 have 896 experts. Imagine you have company with 896 employees. Each question involves 16 employees to figure out what to do in what area of your business. Like a brainstorm to solve problem. Now if you know exactly what you want and in which area you actually need only 1-5 people. Like an agile team instead of all those people that you have. So you can hire just couple Kimi K3 experts. 16 experts are 100B parameters so roughly 1 expert in frontier open source model is 6B parameters. 5 experts is 30B parameters. You can run 27B Qwen 3.6 quantized into int4 on your computer like other people are doing right now.</p>
<p dir="auto">I posted link below to example where they fine tuned model ( take it like a employee training ) for specific task.</p>
]]></description><link>https://citiverse.it/post/https://lemmy.world/comment/25030070</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.world/comment/25030070</guid><dc:creator><![CDATA[vane@lemmy.world]]></dc:creator><pubDate>Thu, 30 Jul 2026 02:46:49 GMT</pubDate></item><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Wed, 29 Jul 2026 21:34:05 GMT]]></title><description><![CDATA[<p dir="auto">How do you get into this? Any articles you can share? OP says you can run on your desktop... How? Doesn't this stuff require huge data centers?</p>
]]></description><link>https://citiverse.it/post/https://lemmy.world/comment/25026374</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.world/comment/25026374</guid><dc:creator><![CDATA[madcaesar@lemmy.world]]></dc:creator><pubDate>Wed, 29 Jul 2026 21:34:05 GMT</pubDate></item><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Wed, 29 Jul 2026 18:29:57 GMT]]></title><description><![CDATA[<p dir="auto">Sure, a benchmark doesn't capture all the subtleties and different use cases, but it does give a general idea of the capabilities of a model. Obviously, you have to run the model and see if it does what you need. But the chart isn't really about the nuance, it's showing how drastically the efficiency of the models has improved in just a year. The fact that we can even reasonably compare a model you can run on a desktop to one that needed a data center just a year ago is phenomenal.</p>
]]></description><link>https://citiverse.it/post/https://lemmy.ml/comment/26966310</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.ml/comment/26966310</guid><dc:creator><![CDATA[yogthos@lemmy.ml]]></dc:creator><pubDate>Wed, 29 Jul 2026 18:29:57 GMT</pubDate></item><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Wed, 29 Jul 2026 17:18:35 GMT]]></title><description><![CDATA[<p dir="auto">according to performance on standard benchmark. somewhat covered by the controversy surrounding the term: benchmaxing.<br />
if you see all benefit as a linear one dimensional height on a bar graph..<br />
its almost like you assume that the previous model gave the same exact answer(same style) and the new mode gave the same exact answer PLUS additional useful information.<br />
it might be convenient if measuring progress was so simple. but unfortunately/fortunately , it is not so simple .<br />
the most important benchmark are the comparison of outcomes on the problems that YOU have &amp; prompts that YOU can(will) write. nothing else matters for YOU.</p>
<ul>
<li>i admit benchmarks are well designed to objectively measure competence on challenging problems that require skill and really need only ONE correct answer.</li>
</ul>
]]></description><link>https://citiverse.it/post/https://lemmy.ml/comment/26965129</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.ml/comment/26965129</guid><dc:creator><![CDATA[leanleft@lemmy.ml]]></dc:creator><pubDate>Wed, 29 Jul 2026 17:18:35 GMT</pubDate></item><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Wed, 29 Jul 2026 17:16:31 GMT]]></title><description><![CDATA[<p dir="auto">This seems like a great time to mention what I built and hosted months ago. <a href="https://masland.tech/ai-efficiency-index/" target="_blank" rel="noopener noreferrer nofollow ugc">https://masland.tech/ai-efficiency-index/</a></p>
<p dir="auto">The AI Efficiency Index tracks the cost/intelligence mix. Low intelligence is useless. Capability scales exponentially. Expensive ≠ better value. MiMo-V2.5 is the leader.</p>
<p dir="auto"><img src="https://lemmy.world/pictrs/image/634318f8-e535-4e9f-ad4a-72c3131ff3fa.png" alt="" class=" img-fluid img-markdown" /></p>
<p dir="auto">If you have any questions feel free to ask. This uses the Artificial Analysis leaderboard. They have "cost per task" which I think is weird.</p>
]]></description><link>https://citiverse.it/post/https://lemmy.world/comment/25022488</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.world/comment/25022488</guid><dc:creator><![CDATA[jaykrown@lemmy.world]]></dc:creator><pubDate>Wed, 29 Jul 2026 17:16:31 GMT</pubDate></item><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Wed, 29 Jul 2026 17:10:05 GMT]]></title><description><![CDATA[<p dir="auto">You can't right now because there is no dataset available for fine tuning. I'm just saying that it's possible to fine tune 8B-35B parameter model in int4 that will outperform those models ex. for single programming language and developer specific problems. You can read example of fine tuned 8B model here <a href="https://fermisense.com/when-machines-take-the-wheel/" target="_blank" rel="noopener noreferrer nofollow ugc">https://fermisense.com/when-machines-take-the-wheel/</a></p>
]]></description><link>https://citiverse.it/post/https://lemmy.world/comment/25022324</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.world/comment/25022324</guid><dc:creator><![CDATA[vane@lemmy.world]]></dc:creator><pubDate>Wed, 29 Jul 2026 17:10:05 GMT</pubDate></item><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Wed, 29 Jul 2026 16:29:52 GMT]]></title><description><![CDATA[<p dir="auto">I'm fairly optimistic that people will figure out how to optimize the models a lot further going forward. One obvious path is to try and separate the reasoning network from the trivia that gets baked into the model, and some work is being done in this area. If you could have a context free reasoning engine and then feed the facts it needs to know on the fly based on the context you're running it in, then you could likely have a much smaller model that's very capable.</p>
]]></description><link>https://citiverse.it/post/https://lemmy.ml/comment/26964198</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.ml/comment/26964198</guid><dc:creator><![CDATA[yogthos@lemmy.ml]]></dc:creator><pubDate>Wed, 29 Jul 2026 16:29:52 GMT</pubDate></item><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Wed, 29 Jul 2026 16:25:58 GMT]]></title><description><![CDATA[<p dir="auto">I run Qwen 3.6 27B quantized down to Q4_K_M or even a variant Q5_K_S on my 8gb VRAM entry level AMD GPU, with support of my CPU and 32gb system RAM. Yes, I also limit the Context Length heavily to something like 18k. It's slow. But the point is, you don't need necessarily 16gb VRAM.</p>
<p dir="auto">But it's better to use faster models for this type of hardware anyway. The MoE type of models (such as 26B A4B, or 35B A3B) are vastly, vastly faster for normal usage, but they are a bit worse in some cases. Also Googles QAT trained model versions also have less RAM requirements without losing much quality.</p>
<p dir="auto">I'm just saying that, so others are not discouraged too much. It's not the same experience as the full version off course, but you can use them with some tricks on weak hardware too.</p>
]]></description><link>https://citiverse.it/post/https://lemmy.ml/comment/26964125</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.ml/comment/26964125</guid><dc:creator><![CDATA[thingsiplay@lemmy.ml]]></dc:creator><pubDate>Wed, 29 Jul 2026 16:25:58 GMT</pubDate></item><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Wed, 29 Jul 2026 15:40:04 GMT]]></title><description><![CDATA[<p dir="auto">Not sure what Moore's law has to do with anything here to be honest. The models you can run locally on a consumer desktop can do real work, and their resource usage is no different from any other software like games that you'd run.</p>
]]></description><link>https://citiverse.it/post/https://lemmy.ml/comment/26963175</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.ml/comment/26963175</guid><dc:creator><![CDATA[yogthos@lemmy.ml]]></dc:creator><pubDate>Wed, 29 Jul 2026 15:40:04 GMT</pubDate></item><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Wed, 29 Jul 2026 14:44:16 GMT]]></title><description><![CDATA[<p dir="auto">How can you run a local model for better coding support than fable/opus?</p>
]]></description><link>https://citiverse.it/post/https://lemmy.dbzer0.com/comment/27332068</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.dbzer0.com/comment/27332068</guid><dc:creator><![CDATA[dnb@lemmy.dbzer0.com]]></dc:creator><pubDate>Wed, 29 Jul 2026 14:44:16 GMT</pubDate></item><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Wed, 29 Jul 2026 14:50:30 GMT]]></title><description><![CDATA[<p dir="auto">Sure they are. Especially with moores law dead. And especially since 2020; the more million dollars behind it, the more they run like one from a inexperienced single-dev on itch.</p>
]]></description><link>https://citiverse.it/post/https://lemmy.zip/comment/27919829</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.zip/comment/27919829</guid><dc:creator><![CDATA[monkdervierte@lemmy.zip]]></dc:creator><pubDate>Wed, 29 Jul 2026 14:50:30 GMT</pubDate></item><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Wed, 29 Jul 2026 14:01:57 GMT]]></title><description><![CDATA[<p dir="auto">I'll give LM studio a go, thanks.</p>
]]></description><link>https://citiverse.it/post/https://programming.dev/comment/25156484</link><guid isPermaLink="true">https://citiverse.it/post/https://programming.dev/comment/25156484</guid><dc:creator><![CDATA[camerondev@programming.dev]]></dc:creator><pubDate>Wed, 29 Jul 2026 14:01:57 GMT</pubDate></item><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Wed, 29 Jul 2026 13:10:17 GMT]]></title><description><![CDATA[<p dir="auto">Ok, thanks! Thaf sounds quite advanced, but I'll have a read afterwards <img src="https://citiverse.it/assets/plugins/nodebb-plugin-emoji/emoji/android/1f642.png?v=b09b169718a" class="not-responsive emoji emoji-android emoji--slightly_smiling_face" style="height:23px;width:auto;vertical-align:middle" title=":)" alt="🙂" /></p>
]]></description><link>https://citiverse.it/post/https://lemmy.world/comment/25017994</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.world/comment/25017994</guid><dc:creator><![CDATA[stuner@lemmy.world]]></dc:creator><pubDate>Wed, 29 Jul 2026 13:10:17 GMT</pubDate></item><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Wed, 29 Jul 2026 13:04:41 GMT]]></title><description><![CDATA[<p dir="auto">I run it using LM Studio, which defaults to Q4 quantization, I think. I was able to put about 10 layers on the GPU with 64k token context. That put me at about 9.1 GB VRAM usage, leaving some room for Video playback xD</p>
]]></description><link>https://citiverse.it/post/https://lemmy.world/comment/25017930</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.world/comment/25017930</guid><dc:creator><![CDATA[stuner@lemmy.world]]></dc:creator><pubDate>Wed, 29 Jul 2026 13:04:41 GMT</pubDate></item><item><title><![CDATA[Reply to GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B that you can run on your desktop on Wed, 29 Jul 2026 12:53:15 GMT]]></title><description><![CDATA[<p dir="auto">You need a GPU with around 16gb vram at a minimum to run qunatized version.</p>
]]></description><link>https://citiverse.it/post/https://lemmy.ml/comment/26960232</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.ml/comment/26960232</guid><dc:creator><![CDATA[yogthos@lemmy.ml]]></dc:creator><pubDate>Wed, 29 Jul 2026 12:53:15 GMT</pubDate></item></channel></rss>