Z.AI Gives the Weights Away. Tokens Are Now 86.5% of Revenue.
TL;DR
On August 31 Z.AI, the Chinese lab known domestically as Zhipu that publishes its GLM weights under MIT, filed its first half-year report as a listed company. Revenue for the six months to June 30 was RMB 953.9 million, roughly US$142 million, up 399.7% year on year. The shape of that number is the story: 86.5% of it came from the open platform and API line, against 15.2% a year earlier, while the segment where enterprises license a general-purpose model and run it themselves fell 54.6% to RMB 67 million. Gross profit was RMB 251.6 million, a 26.4% margin. R&D was RMB 2.13 billion, more than twice revenue. Net loss was RMB 2.07 billion. Bloomberg reported the top line landed about 30% below the average analyst estimate.
The mix flipped in twelve months
Five days ago the interesting Z.AI fact was that GLM-5.3-Flash shipped as a 320B mixture-of-experts model with MIT weights at $0.15 in and $0.50 out per million tokens. This filing is the other half of that fact: the profit and loss statement underneath the giveaway.
A year ago Z.AI was mostly a licensing company. Enterprises bought a general-purpose model, deployed it inside their own perimeter, and paid for the privilege. That line was the majority of the business. Today it is 7% of revenue and shrinking at 54.6% a year, and almost everything the company earns arrives as metered tokens.
Publishing the weights is the best-funded marketing campaign in the industry. Anyone can take the recipe home, and it turns out most people would rather someone else run the kitchen. The MaaS platform now reports more than 7.4 million enterprise and developer accounts, and every one of them that gives up on provisioning H-series cards becomes an API line item.
Z.AI's own chief executive called this a quarter early. Discussing the 2025 annual results in March, Zhang Peng said clients who first tried to deploy the open-source models locally were "gradually shifting" toward the cloud API, at least in part, according to the South China Morning Post. The H1 segment split is what that sentence looks like once it hits the books.
Twenty-six cents on the yuan
Here is the part that should give you pause if your product sits on cheap GLM tokens. Revenue grew 399.7%. Gross profit grew 163.7%. When the top line runs at more than twice the speed of gross profit, margin is being spent to buy it.
The FY2024 and FY2025 margins come from the annual revenue and gross profit lines on the 2513.HK financial record: RMB 175.9 million on RMB 312.4 million, then RMB 296.7 million on RMB 724.3 million. The H1 2026 figure is straight from the interim numbers. Nobody had to model anything.
What eats the difference is serving cost. Inference is not software margin, it is rented silicon plus power, and the price per token is set by whoever is willing to lose the most money this quarter. That fight is currently loud in China, which is why Bloomberg filed the miss under an AI price war rather than under weak demand.
For a live example, check the Z.AI price sheet. GLM-5.3-Flash lists at $0.15 in and $0.50 out per million tokens, and is currently running a 50% promotional discount through September 9. A company that just reported a 26.4% gross margin is selling its volume model at half price. Draw your own conclusions about the H2 margin line.
R&D costs more than the entire business
The other number worth sitting with: research and development was RMB 2.13 billion for the half, up 33.6%. That is 2.2 times the revenue it produced. Net loss was RMB 2.07 billion, which narrowed 12.1% year on year, while adjusted net loss widened 12.1% to RMB 1.96 billion. The two moving in opposite directions is a reminder to read which loss line a headline is quoting.
Scale check: this half alone brought in more revenue than the whole of 2025, when Z.AI booked RMB 724.3 million. The business is genuinely compounding. It is just compounding into a cost base that is compounding faster.
Why the private-deployment line fell off
The 54.6% drop in enterprise general-purpose model revenue is the most instructive line in the filing, and it is not obviously a failure. It is partly cannibalisation the company chose.
If you are a bank in 2024 and you want a capable model behind your firewall, you buy a licensed deployment and a support contract. If you are the same bank in 2026, the weights for a frontier-adjacent GLM model are sitting on Hugging Face under MIT, so the licence is worth nothing and the only thing left to sell you is the operational burden of running it. Most buyers look at the GPU quote, the ops headcount, and the upgrade treadmill, and pick the API instead.
That is a good trade for Z.AI on volume and a bad trade on margin quality. Licensing revenue is high-margin and lumpy. Token revenue is low-margin, recurring, and priced by the most desperate competitor in the market. The company swapped one for the other in about four quarters.
What this changes if you build on GLM
- Your cheap tokens are being subsidised. A 26.4% gross margin and a 50% promotional discount are not a stable price floor. Budget for list price, not promo price, and re-check the price sheet before you commit a workload.
- The weights are the hedge, and you should actually hold them. The MIT licence is the reason a price move is an inconvenience rather than an outage. Pull the checkpoint you depend on now and keep a copy. That is the whole point of open weights and almost nobody does it.
- Watch the segment mix, not the growth rate. 399.7% growth reads great. The same filing says the durable, high-margin revenue line halved. Growth composed entirely of the low-margin segment is a different company than growth spread across both.
- Vendor solvency is a dependency. R&D at 2.2 times revenue is funded by a war chest, not by operations. If you are routing production traffic to one API, know how long that war chest lasts, or keep a second provider warm.
The caveats
These are interim figures for a company two quarters into its Hong Kong listing, and interim reports are unaudited. The FY2024 and FY2025 gross margins here are computed from reported revenue and gross profit, not lifted from a margin line in the filing. Loss definitions vary between the statutory number and the adjusted number, and different outlets quote different ones, which is how the same company can be described as narrowing and widening its loss in the same paragraph.
Yuan to dollar conversion follows the figure the wires used, about US$142 million for RMB 953.9 million. Analyst consensus is Bloomberg's, not the company's, and the roughly 30% shortfall is against that consensus rather than against any guidance Z.AI published.
Key Takeaways
- Z.AI reported RMB 953.9 million of H1 2026 revenue, up 399.7%, with 86.5% of it from the open platform and API line against 15.2% a year earlier.
- Enterprise general-purpose large model revenue, the private-deployment licensing business, fell 54.6% to RMB 67 million and is now 7% of the company.
- Gross margin has gone from 56.3% in FY2024 to 41.0% in FY2025 to 26.4% in H1 2026 as token pricing compressed.
- R&D of RMB 2.13 billion and a net loss of RMB 2.07 billion each run at roughly 2.2 times half-year revenue.
- Open weights are working exactly as designed as a distribution channel, and the resulting business is an inference business with inference margins.
- If your product depends on GLM API pricing, the MIT weights are your only real hedge. Download them before you need them.
Sources: South China Morning Post on the H1 2026 results, Investing.com results summary, Bloomberg, South China Morning Post on the FY2025 annual report, Z.AI (2513.HK) financials, Z.AI API pricing, Z.AI GLM-5.3-Flash announcement, Z.AI on Hugging Face