Skip to content

June 18, 2026

Redesigned model registration, FP8/FP4 quantization support, and clearer billing indicators during deployment.

  • Model registration redesigned — registering a Hugging Face repo now auto-detects the model’s capabilities, and every derived value can still be overridden.
  • FP8 / FP4 (NVFP4) quantized models are supported with hardware-aware GPU matching — GPUs that can’t run a model’s quantization are excluded from candidates, with the reason shown.
  • The price tag on a starting deployment now switches to a real $/hr figure exactly when billing becomes active, instead of showing a placeholder throughout the whole startup process.
  • Continued reliability improvements for multi-GPU deployments, including better handling of larger models split across several GPUs.