Overview
Morgan Stanley analysts assess that the emergence of lower‑cost open‑weight generative AI (GenAI) models could compress token pricing to roughly $1.75 per million tokens and increase competition among model laboratories, thereby pressuring their profit margins.
Return on Invested Capital Estimates
The bank estimates that model providers operating one gigawatt of Nvidia GB300 infrastructure could generate ROIC in the range of 20 % to 60 %, assuming token throughput of 2,000 to 3,500 tokens per second per GPU. For hyperscalers that rent out GPU capacity, the projected ROIC lies between 23 % and 39 %, depending on hourly pricing.
Drivers Supporting Attractive Economics
1. Computing power remains scarce, and enterprises continue to rely on cloud providers for open‑weight inference workloads.
2. Lower‑cost models are expected to drive substantially higher usage volumes, allowing hyperscalers to adjust capacity prices based on supply‑demand dynamics, which can boost total profit growth more than percentage margins.
3. Improvements in token throughput are being realized through newer chips, faster interconnects, more efficient model architectures, and request‑batching software. Amazon (NASDAQ:AMZN) and Alphabet (NASDAQ:GOOGL) employ proprietary silicon such as Trainium and Tensor Processing Units to lower compute costs.
4. Open‑weight workloads can generate ancillary revenue from managed APIs, GPU rentals, databases, storage, and security tools; low‑priced model access may act as a loss leader that fuels higher‑margin cloud spending.
Analyst Recommendations
Morgan Stanley maintains Overweight ratings on Amazon and Alphabet, highlighting their ability to monetize scarce computing capacity at reduced cost. The analysts also advise monitoring Meta Platforms’ (NASDAQ:META) Muse models and Alphabet’s Gemini Flash products for early signs of pricing pressure, adoption trends, and throughput improvements.