June 18, 2026
Redesigned model registration, FP8/FP4 quantization support, and clearer billing indicators during deployment.
Improvements
Section titled “Improvements”- Model registration redesigned — registering a Hugging Face repo now auto-detects the model’s capabilities, and every derived value can still be overridden.
- FP8 / FP4 (NVFP4) quantized models are supported with hardware-aware GPU matching — GPUs that can’t run a model’s quantization are excluded from candidates, with the reason shown.
- The price tag on a starting deployment now switches to a real
$/hrfigure exactly when billing becomes active, instead of showing a placeholder throughout the whole startup process. - Continued reliability improvements for multi-GPU deployments, including better handling of larger models split across several GPUs.