
The GPU Rental Bubble Burst. Here's What That Means for AI Prices.
Don't buy H100s.
Eugene Cheah spent the last year analyzing GPU economics. He cofounded Featherless.AI, an inference platform, so he's been watching this closely. His conclusion: buying H100s right now is a bad investment.
The numbers back him up. In 2023, an H100 rented for $8+ per hour if you could get one at all. By August 2024, auction platforms had them at $1-2 per hour. That's a 75% collapse in 18 months.
GPU rental prices are a leading indicator. They're the floor underneath AI subscription prices. When the floor drops, eventually everything above it has to follow.
The $50,000 math problem
An H100 SXM GPU costs about $50,000 to deploy in a data center. That's hardware, setup, networking, all the capex before you spend a dime on electricity or cooling.
NVIDIA's 2023 investor pitch: rent these at $4 per hour, get your money back in under 2 years, then collect over $100,000 per GPU per year in pure profit. Data center operators believed it. Foundation model companies believed it. Investors poured in tens of billions. GPUs were the new gold rush.
The demand forecast was wrong.
What killed demand
Open weights got good enough. When Meta's Llama 3 and models like DeepSeek V2 reached GPT-4 class performance, the economics of training from scratch fell apart. Why spend millions when you can fine-tune an existing model for a fraction of the cost? Fine-tuning might need one GPU node. Training needs 16+. That demand disappeared.
At the same time, the foundation model gold rush ended. By Eugene's count, fewer than 50 teams worldwide actually need 16+ H100 nodes for foundation model training. There are more than 50 such clusters available. Supply exceeds demand. If you can't beat Llama 3 and you don't have a fundamentally new approach, investors aren't interested.
Then the prepaid capacity hit the resale market. Those 3-5 year contracts with 50-100% upfront payment that data centers pushed in 2023? The companies that signed them finished training, pivoted to fine-tuning, or went under. Now they're stuck paying for GPUs they don't need, so they're reselling through RunPod, Vast.ai, Together.ai. Since the hardware is already paid for, any revenue is better than nothing. They undercut everyone.
The break-even math
Eugene's analysis shows three tiers:
- Above $2.85/hour: You beat stock market returns, assuming 100% utilization which never happens
- Below $2.85/hour: You'd make more money in an index fund
- Below $1.65/hour: You're losing money over the hardware's lifespan
Current market rates on auction platforms: $1-2 per hour. GPU owners who bought in 2023-2024 are looking at ugly returns. Prices keep dropping roughly 40% per year. NVIDIA's projection of $4/hour for 4 years evaporated in 18 months.
Inference doesn't need H100s
Most AI workloads aren't training. They're inference, running models that have already been trained. For that, you don't need what makes H100s expensive.
NVIDIA's own L40S offers one third the performance at one fifth the price. It can't do multi-node training but works fine for inference. AMD's MI300X and Intel's Gaudi 3 have more memory and compute than an H100 at lower prices. They're unproven for massive training clusters but work as drop-in replacements for single-node inference.
Then there's the flood of ex-crypto GPUs. With Ethereum's move to proof-of-stake, consumer GPUs that used to mine are now available for inference. They won't train anything but for running models under 10 billion parameters, they're dirt cheap.
What this means for AI pricing
The GPU oversupply isn't going away. Blackwell and H200 chips are coming online, adding more supply. Prices will keep falling.
When it costs less to run models, it should cost less to use them. But subscription prices are sticky. ChatGPT Plus and Claude Pro are both still $20 per month. The providers are capturing the efficiency gains as margin.
Open weights set a ceiling. When Llama 3 405B and DeepSeek V2 hit GPT-4 quality, they created a free alternative to commercial APIs. If the free option is 90% as good, the paid option needs to justify the remaining 10%.
Aggregators like Admix bundle provider access behind a monthly credit allowance. Starter is $10 a month, or $8 a month billed annually, and unlocks the 350+ model catalog.
The upside
Cheaper GPUs are good for everyone except the people who bought H100 clusters at the peak. The cost barrier to AI experimentation is disappearing. You can fine-tune models, run inference, and build AI products for a fraction of what it cost a year ago.
For users, the math is simple. If you're spending $40-60 per month across multiple AI subscriptions, consolidating through an aggregator costs a fraction of that. Compute is becoming a commodity. The question is whether subscription prices will reflect that reality or whether aggregators will force the issue.
FAQ
Should I buy GPUs for AI work?
Probably not. The rental market is oversupplied, meaning you can rent at or below cost. Buying locks you into a depreciating asset. Rent when you need it.
Will AI subscriptions get cheaper?
Eventually. Compute costs have dropped 80-90% but subscription prices haven't budged. Aggregators like Admix already offer lower prices. The major providers will have to follow or lose market share.
Are H100 prices still falling?
Yes. The market shows roughly 40% annual declines, and new hardware (H200, Blackwell) adds supply. This trend continues through 2026-2027.
What's an "open weights" model?
Models like Llama 3, Mistral, and DeepSeek that distribute their parameters freely. Anyone can download and run them. They're not technically "open source" due to licensing restrictions, but they're freely available and put pressure on commercial API pricing.
Try all the models mentioned in this article
Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.
Start free on Admix