Z.ai Served GLM-5.3-Flash Entirely on Chinese AI Chips
-
This post did not contain any content.
Z.ai Served GLM-5.3-Flash Entirely on Chinese AI Chips
Z.ai says the anonymous GLM-5.3-Flash preview ran across tens of thousands of domestic Chinese accelerators at per-token cost comparable to Nvidia GPUs. The company named no chip vendor and published no throughput or power figures, and none of the serving results has been independently audited.
Implicator.ai (www.implicator.ai)
-
This post did not contain any content.
Z.ai Served GLM-5.3-Flash Entirely on Chinese AI Chips
Z.ai says the anonymous GLM-5.3-Flash preview ran across tens of thousands of domestic Chinese accelerators at per-token cost comparable to Nvidia GPUs. The company named no chip vendor and published no throughput or power figures, and none of the serving results has been independently audited.
Implicator.ai (www.implicator.ai)
What is this headline?
-
This post did not contain any content.
Z.ai Served GLM-5.3-Flash Entirely on Chinese AI Chips
Z.ai says the anonymous GLM-5.3-Flash preview ran across tens of thousands of domestic Chinese accelerators at per-token cost comparable to Nvidia GPUs. The company named no chip vendor and published no throughput or power figures, and none of the serving results has been independently audited.
Implicator.ai (www.implicator.ai)
Ox Alpha Stealth Model Launches With 100T Token Capacity and GLM 5.3 Fingerprints | HuggingNews
An anonymous multimodal AI model called Ox Alpha has launched for free testing on OpenRouter and OpenCode with a 1 million token context window and a daily cap…
HuggingNews (huggingnews.com)
Assuming they are related; A capacity @ 100 Trillion tokens a day, is something of a statement ! According to this site that is 1/4 of all global current token generation, which supposedly are at 390T tokens a day.
-
What is this headline?
That China is embargoed and is supposed to have no access to that types of chips.
Recent months showed a huge deal of ingenuity achieve almost competitive chips. Parts of their domestic use seems to he covered, already.
-
Ox Alpha Stealth Model Launches With 100T Token Capacity and GLM 5.3 Fingerprints | HuggingNews
An anonymous multimodal AI model called Ox Alpha has launched for free testing on OpenRouter and OpenCode with a 1 million token context window and a daily cap…
HuggingNews (huggingnews.com)
Assuming they are related; A capacity @ 100 Trillion tokens a day, is something of a statement ! According to this site that is 1/4 of all global current token generation, which supposedly are at 390T tokens a day.
Wait can tokens just be directly compared like that? My impression is that token cost can vary by 2 orders of magnitude, depending on model, because the actual work of computation varies by that much between models.
-
Wait can tokens just be directly compared like that? My impression is that token cost can vary by 2 orders of magnitude, depending on model, because the actual work of computation varies by that much between models.
Yeah this is a bad metric. GLM outputs a lot of reasoning tokens.

Ciao! Sembra che tu sia interessato a questa conversazione, ma non hai ancora un account.
Stanco di dover scorrere gli stessi post a ogni visita? Quando registri un account, tornerai sempre esattamente dove eri rimasto e potrai scegliere di essere avvisato delle nuove risposte (tramite email o notifica push). Potrai anche salvare segnalibri e votare i post per mostrare il tuo apprezzamento agli altri membri della comunità.
Con il tuo contributo, questo post potrebbe essere ancora migliore 💗
Registrati Accedi
Citiverse è un progetto che si basa su NodeBB ed è federato! | Categorie federate | Chat | 📱 Installa web app o APK | 🧡 Donazioni | Privacy Policy