Skip to content

April 12, 2026

Serverless Inference v1 launches — deploy models behind an API in a few clicks, with an in-console playground.

Launched Serverless Inference — deploy a model behind an API endpoint without managing GPUs yourself. The Models page was redesigned around it: a table view with official and custom model tabs, an in-browser Playground, and Hugging Face token management for private models. Capacity scales automatically with demand.

  • Simplified deployment by replacing the separate deploy page with an inline flow, cutting down on page navigation.
  • The host-client message thread is now easier to follow, with proper conversation grouping.