Skip to content

May 3, 2026

A new per-GPU-second pricing model for Serverless with a cost estimate preview, and better multi-turn chat performance.

Serverless moved to per-GPU-second pricing with benchmark-based suggested pricing, and the Playground now shows an estimated cost, time, and VRAM requirement before you run a job.

  • Improved chat inference consistency by routing follow-up messages in the same conversation back to the same server, improving cache reuse and response times.
  • Reduced deployment failures (timeouts, out-of-memory) for large models running on external buffer capacity.