<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[If this is true, the hyperscalers are toast]]></title><description><![CDATA[<em>This post did not contain any content.</em>

<div class="row mt-3"><div class="card col-md-9 col-lg-6 position-relative link-preview p-0">



<a href="https://klementoninvesting.substack.com/p/if-this-is-true-the-hyperscalers" title="If this is true, the hyperscalers are toast">
<img src="https://substackcdn.com/image/fetch/$s_!RLl7!,w_1200,h_675,c_fill,f_jpg,q_auto:good,fl_progressive:steep,g_auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4827262-1cea-49a3-ae80-6d2267663426_640x792.jpeg" class="card-img-top not-responsive" style="max-height: 15rem;" alt="Link Preview Image" onerror="this.parentElement.remove()" />
</a>



<div class="card-body">
<h5 class="card-title">
<a class="text-decoration-none" href="https://klementoninvesting.substack.com/p/if-this-is-true-the-hyperscalers">
If this is true, the hyperscalers are toast
</a>
</h5>
<p class="card-text line-clamp-3">In my regular research (behind a paywall), I have been saying for a while that I think the future of AI is not large language models (LLM), but small language models (SLM) run on local desktop computers or even mobile phones.</p>
</div>
<a href="https://klementoninvesting.substack.com/p/if-this-is-true-the-hyperscalers" class="card-footer text-body-secondary small d-flex gap-2 align-items-center lh-2">



<img src="https://substackcdn.com/image/fetch/$s_!PuaX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc774836-b649-4979-b33f-a3984c09d41e%2Ffavicon-16x16.png" alt="favicon" class="not-responsive overflow-hiddden" style="max-width: 21px; max-height: 21px;" onerror="this.remove()"/>































<p class="d-inline-block text-truncate mb-0"> <span class="text-secondary">(klementoninvesting.substack.com)</span></p>
</a>
</div></div>]]></description><link>https://citiverse.it/topic/5c1ed8d1-28d1-43e0-8795-fba10239c397/if-this-is-true-the-hyperscalers-are-toast</link><generator>RSS for Node</generator><lastBuildDate>Sun, 23 Aug 2026 15:34:13 GMT</lastBuildDate><atom:link href="https://citiverse.it/topic/5c1ed8d1-28d1-43e0-8795-fba10239c397.rss" rel="self" type="application/rss+xml"/><pubDate>Thu, 20 Aug 2026 13:04:10 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to If this is true, the hyperscalers are toast on Fri, 21 Aug 2026 17:12:26 GMT]]></title><description><![CDATA[<p dir="auto">Possibly, I haven't looked at how easy it is to get your hands on one of those.</p>
]]></description><link>https://citiverse.it/post/https://lemmy.ml/comment/27373759</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.ml/comment/27373759</guid><dc:creator><![CDATA[yogthos@lemmy.ml]]></dc:creator><pubDate>Fri, 21 Aug 2026 17:12:26 GMT</pubDate></item><item><title><![CDATA[Reply to If this is true, the hyperscalers are toast on Fri, 21 Aug 2026 13:48:58 GMT]]></title><description><![CDATA[<p dir="auto">What about those small prebuilt workstations with Strix/Gorgon Halo chips and big ram pools? Chinese OEMs like Beeline and GMKTec are offering some sleek little boxes. I guess it still technically not unified on a system level, but you allot most of it to VRAM in the BIOS regardless. You can probably get comparable performance with a fraction of the cost.</p>
]]></description><link>https://citiverse.it/post/https://lemmy.ml/comment/27370133</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.ml/comment/27370133</guid><dc:creator><![CDATA[shinkantrain@lemmy.ml]]></dc:creator><pubDate>Fri, 21 Aug 2026 13:48:58 GMT</pubDate></item><item><title><![CDATA[Reply to If this is true, the hyperscalers are toast on Fri, 21 Aug 2026 13:19:17 GMT]]></title><description><![CDATA[<p dir="auto">If you absolutely need one now, a mac is probably the cheapest way to run them because of the unified memory. With any x86 solution you have to get a separate video card with at least 32gb vram to run a decent local model. However, if you wait a bit then you can probably get a dedicated chip a lot cheaper in the near future <a href="https://wccftech.com/alibabas-tsmc-built-5nm-risc-v-chip-xuantie-c950-now-runs-qwen-3-8-27b-model-natively-unlocking-massive-vertical-integration-tailwinds/" target="_blank" rel="noopener noreferrer nofollow ugc">https://wccftech.com/alibabas-tsmc-built-5nm-risc-v-chip-xuantie-c950-now-runs-qwen-3-8-27b-model-natively-unlocking-massive-vertical-integration-tailwinds/</a></p>
]]></description><link>https://citiverse.it/post/https://lemmy.ml/comment/27369750</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.ml/comment/27369750</guid><dc:creator><![CDATA[yogthos@lemmy.ml]]></dc:creator><pubDate>Fri, 21 Aug 2026 13:19:17 GMT</pubDate></item><item><title><![CDATA[Reply to If this is true, the hyperscalers are toast on Fri, 21 Aug 2026 13:15:27 GMT]]></title><description><![CDATA[<p dir="auto">I expect doing ASICs for models will work even if they keep improving. It'll be like regular chips getting new versions. You buy a chip with a specific model etched into it, and if it does what you need great. Next year, a new version comes out. So, it's actually a feature since it allows companies to keep selling new chips.</p>
<p dir="auto">It does look like we are entering diminishing returns territory though. The biggest evidence for this is that Chinese companies have now basically caught up to Anthropic and OpenAI. If the progress at the frontier was still happening at the same rate, then the gap wouldn't be closing so quickly. There's also a lot less noticeable difference between stuff like Claude 4.6 and Claude 5. When they went from 3.x to 4.x it was very noticeable. And at least for agentic coding, most of the improvement seems to come from the harness now. I expect improvements will continue, but at a much more gradual pace. It's also possible people will figure out a new architecture that's superior to LLMs, or works with them. World models are one promising area already being explored.</p>
]]></description><link>https://citiverse.it/post/https://lemmy.ml/comment/27369680</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.ml/comment/27369680</guid><dc:creator><![CDATA[yogthos@lemmy.ml]]></dc:creator><pubDate>Fri, 21 Aug 2026 13:15:27 GMT</pubDate></item><item><title><![CDATA[Reply to If this is true, the hyperscalers are toast on Fri, 21 Aug 2026 14:48:00 GMT]]></title><description><![CDATA[<p dir="auto">Yes. The issue is not that ARM systems don’t/can’t communicate with a dGPU over PCIe. It’s that x86 systems (edit: running conventional x86 desktop OSes) have to basically pretend they’re using PCIe to communicate with an iGPU. (Not literally, but a lot of the primitives they use were inherited from AGP/PCI and don’t make a lot of sense for an SoC.)</p>
]]></description><link>https://citiverse.it/post/https://midwest.social/comment/25718087</link><guid isPermaLink="true">https://citiverse.it/post/https://midwest.social/comment/25718087</guid><dc:creator><![CDATA[kibiz0r@midwest.social]]></dc:creator><pubDate>Fri, 21 Aug 2026 14:48:00 GMT</pubDate></item><item><title><![CDATA[Reply to If this is true, the hyperscalers are toast on Fri, 21 Aug 2026 11:27:10 GMT]]></title><description><![CDATA[<p dir="auto">One of my first posts here was asking how to make use of local LLM the best. People said to buy the most beefy laptop within my budget.</p>
<p dir="auto">Is that really the most reasonable take?</p>
]]></description><link>https://citiverse.it/post/https://lemmy.ml/comment/27368113</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.ml/comment/27368113</guid><dc:creator><![CDATA[socialistvibes01@lemmy.ml]]></dc:creator><pubDate>Fri, 21 Aug 2026 11:27:10 GMT</pubDate></item><item><title><![CDATA[Reply to If this is true, the hyperscalers are toast on Fri, 21 Aug 2026 10:21:10 GMT]]></title><description><![CDATA[<p dir="auto">Aren't there arm workstations that still take pcie cards and have a similar architecture to x86?</p>
]]></description><link>https://citiverse.it/post/https://sh.itjust.works/comment/27010153</link><guid isPermaLink="true">https://citiverse.it/post/https://sh.itjust.works/comment/27010153</guid><dc:creator><![CDATA[jumping_redditor@sh.itjust.works]]></dc:creator><pubDate>Fri, 21 Aug 2026 10:21:10 GMT</pubDate></item><item><title><![CDATA[Reply to If this is true, the hyperscalers are toast on Fri, 21 Aug 2026 09:03:23 GMT]]></title><description><![CDATA[<p dir="auto">I wouldn't say it's totally legacy. (v)ram bandwidth does matter and while the Apple M-series chips do well with their unified memory architecture don't forget it's fixed because it's part of the CPU chip.</p>
]]></description><link>https://citiverse.it/post/https://lemmy.ml/comment/27366790</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.ml/comment/27366790</guid><dc:creator><![CDATA[stsquad@lemmy.ml]]></dc:creator><pubDate>Fri, 21 Aug 2026 09:03:23 GMT</pubDate></item><item><title><![CDATA[Reply to If this is true, the hyperscalers are toast on Fri, 21 Aug 2026 13:43:58 GMT]]></title><description><![CDATA[<p dir="auto">I think this all hinges on whether or not progress will slow down. For example I've heard of some companies etching AI models onto silicon, but nobody will buy those if the model is obsolete in a year. But if the models stop improving, then these chips might be worth the investment since they will be way more efficient for inference.</p>
<p dir="auto">So that's one of the biggest questions in AI right now. Are we going to hit a wall? It does kind of seem like the big models aren't improving as much anymore, and the small models are catching up. But at the same time, Moore's law has been going for way longer than people expected, maybe AI will be the same.</p>
<p dir="auto">Edit: wording on last sentence of first paragraph</p>
]]></description><link>https://citiverse.it/post/https://sh.itjust.works/comment/27009308</link><guid isPermaLink="true">https://citiverse.it/post/https://sh.itjust.works/comment/27009308</guid><dc:creator><![CDATA[hirihit640@sh.itjust.works]]></dc:creator><pubDate>Fri, 21 Aug 2026 13:43:58 GMT</pubDate></item><item><title><![CDATA[Reply to If this is true, the hyperscalers are toast on Fri, 21 Aug 2026 03:14:16 GMT]]></title><description><![CDATA[<p dir="auto">Hence “x86 (desktop OS) space”. It’s not intrinsically part of x86, but it has settled in as a conventional piece of x86 desktop OSes. x86 consoles and ARM desktops don’t assume the same principle, and they can get a lot more mileage out of SoCs as a result.</p>
]]></description><link>https://citiverse.it/post/https://midwest.social/comment/25712383</link><guid isPermaLink="true">https://citiverse.it/post/https://midwest.social/comment/25712383</guid><dc:creator><![CDATA[kibiz0r@midwest.social]]></dc:creator><pubDate>Fri, 21 Aug 2026 03:14:16 GMT</pubDate></item><item><title><![CDATA[Reply to If this is true, the hyperscalers are toast on Fri, 21 Aug 2026 01:19:59 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto">he RAM/VRAM divide is kind of a legacy thing at this point, doing more harm than good in the x86</p>
</blockquote>
<p dir="auto">That has nothing to do with x86 though...</p>
]]></description><link>https://citiverse.it/post/https://lemmy.world/comment/25403898</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.world/comment/25403898</guid><dc:creator><![CDATA[ms_lane@lemmy.world]]></dc:creator><pubDate>Fri, 21 Aug 2026 01:19:59 GMT</pubDate></item><item><title><![CDATA[Reply to If this is true, the hyperscalers are toast on Thu, 20 Aug 2026 22:42:30 GMT]]></title><description><![CDATA[<p dir="auto">Alibaba just announced a chip specifically for running local models. We'll see what it ends up going for. <a href="https://wccftech.com/alibabas-tsmc-built-5nm-risc-v-chip-xuantie-c950-now-runs-qwen-3-8-27b-model-natively-unlocking-massive-vertical-integration-tailwinds/" target="_blank" rel="noopener noreferrer nofollow ugc">https://wccftech.com/alibabas-tsmc-built-5nm-risc-v-chip-xuantie-c950-now-runs-qwen-3-8-27b-model-natively-unlocking-massive-vertical-integration-tailwinds/</a></p>
]]></description><link>https://citiverse.it/post/https://lemmy.ml/comment/27360649</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.ml/comment/27360649</guid><dc:creator><![CDATA[yogthos@lemmy.ml]]></dc:creator><pubDate>Thu, 20 Aug 2026 22:42:30 GMT</pubDate></item><item><title><![CDATA[Reply to If this is true, the hyperscalers are toast on Thu, 20 Aug 2026 22:41:38 GMT]]></title><description><![CDATA[<p dir="auto">Training happens once per model, but inference is an ongoing process. So, there's going to be a huge amount of energy saving if we move to using local models.</p>
]]></description><link>https://citiverse.it/post/https://lemmy.ml/comment/27360641</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.ml/comment/27360641</guid><dc:creator><![CDATA[yogthos@lemmy.ml]]></dc:creator><pubDate>Thu, 20 Aug 2026 22:41:38 GMT</pubDate></item><item><title><![CDATA[Reply to If this is true, the hyperscalers are toast on Thu, 20 Aug 2026 19:44:31 GMT]]></title><description><![CDATA[<p dir="auto">SLM's are still trained on massive datacenters as large models and then quantized down. But yes for inference there's hope in the future</p>
]]></description><link>https://citiverse.it/post/https://lemmy.ml/comment/27357852</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.ml/comment/27357852</guid><dc:creator><![CDATA[geneva_convenience@lemmy.ml]]></dc:creator><pubDate>Thu, 20 Aug 2026 19:44:31 GMT</pubDate></item><item><title><![CDATA[Reply to If this is true, the hyperscalers are toast on Thu, 20 Aug 2026 18:53:21 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto">but I suspect they are still going to be bigger than most can comfortably host for time being</p>
</blockquote>
<p dir="auto">ARM and CXMT both have the potential to change this situation rather dramatically.</p>
<p dir="auto">ARM: The RAM/VRAM divide is kind of a legacy thing at this point, doing more harm than good in the x86 (desktop OS) space, but ARM doesn’t have the same baggage.</p>
<p dir="auto">CXMT: We know the DRAM cartel have previously engaged in price-fixing, and the current shortage looks suspiciously similar to their old behavior. When confronted by a new challenger, they might be forced to actually compete.</p>
]]></description><link>https://citiverse.it/post/https://midwest.social/comment/25706715</link><guid isPermaLink="true">https://citiverse.it/post/https://midwest.social/comment/25706715</guid><dc:creator><![CDATA[kibiz0r@midwest.social]]></dc:creator><pubDate>Thu, 20 Aug 2026 18:53:21 GMT</pubDate></item><item><title><![CDATA[Reply to If this is true, the hyperscalers are toast on Thu, 20 Aug 2026 17:53:04 GMT]]></title><description><![CDATA[<p dir="auto">I don't doubt that locally hostable models will be important for agentic tasks but I suspect they are still going to be bigger than most can comfortably host for time being. I still think there will be a place for the super large models for more complex reasoning although how much will be due to the intrinsic knowledge in the weights and how much due to the plumbing around them remains too be seen.</p>
<p dir="auto">And if course LLM's are not going to be the end point of the search for AGI. Whatever their architecture they will still need copious amounts of compute.</p>
]]></description><link>https://citiverse.it/post/https://lemmy.ml/comment/27355946</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.ml/comment/27355946</guid><dc:creator><![CDATA[stsquad@lemmy.ml]]></dc:creator><pubDate>Thu, 20 Aug 2026 17:53:04 GMT</pubDate></item></channel></rss>