May 3, 2026
A new per-GPU-second pricing model for Serverless with a cost estimate preview, and better multi-turn chat performance.
Major features
Section titled “Major features”Estimate before you run
Section titled “Estimate before you run”Serverless moved to per-GPU-second pricing with benchmark-based suggested pricing, and the Playground now shows an estimated cost, time, and VRAM requirement before you run a job.
Improvements
Section titled “Improvements”- Improved chat inference consistency by routing follow-up messages in the same conversation back to the same server, improving cache reuse and response times.
- Reduced deployment failures (timeouts, out-of-memory) for large models running on external buffer capacity.