News
NVIDIA HGX B300 on FPT AI Factory Powers the Next Phase of AI Token Economics
FPT AI Factory is evolving to support this next phase of AI operations. Enhanced with NVIDIA HGX B300 GPUs and tightly integrated with a cloud-native technology stack, the platform is designed to serve as a high-performance engine for inference-heavy and agentic AI workloads.
As AI moves from experimentation to scaled deployment, the economics of inference are becoming a defining factor in enterprise competitiveness. The rise of AI-native applications, reasoning models, and autonomous agents is shifting attention away from managing infrastructure components and toward operating production AI workloads that continuously serve users and generate tokens at scale. In this environment, tokens are no longer just a technical metric. They are becoming a core measure of AI performance, cost, and business value.

The next-gen inference cloud built for AI reasoning
When applications become more agentic, workloads require continuous evaluation, fine-tuning, and orchestration across models and agents in ongoing development loops, contributing to growing token volumes across AI workflows. Token volumes continue to rise across development and production loops, making metrics such as tokens per second, time to first token, and end-to-end latency increasingly critical alongside model quality. The infrastructure behind these workloads must therefore do more than provide raw compute. It must support speed, stability, and economic efficiency throughout the AI lifecycle.
FPT AI Factory is evolving to support this next phase of AI operations. Enhanced with NVIDIA HGX B300 GPUs and tightly integrated with a cloud-native technology stack, the platform is designed to serve as a high-performance engine for inference-heavy and agentic AI workloads. In practical terms, that means giving AI teams the infrastructure they need to generate, process, and scale tokens more efficiently, while reducing the operational burden of managing complex clusters and runtimes.

HGX B300 delivers breakthrough performance on the most complex workloads from training, agentic systems, and reasoning
Built on the NVIDIA Blackwell Ultra architecture, NVIDIA HGX B300 is designed to handle some of the industry’s most demanding AI workloads, from training and reasoning to multimodal systems and agentic applications. With 288GB of memory per GPU and 2.1TB of total GPU memory in a single node, the platform supports trillion-parameter models and long-context reasoning pipelines with fewer constraints related to memory and inter-node communication. NVIDIA HGX B300 also delivers up to 1.5 times higher dense FP4 performance than NVIDIA Blackwell GPUs, giving enterprises greater throughput and efficiency for advanced AI workloads.
These advancements translate directly into AI economics. Organizations running on FPT AI Factory can achieve up to 66% lower inference costs, 49% reduced training costs, and up to 2.95x better cost-per-token optimization. These gains give organizations and AI practitioners greater economic headroom to scale models and serve more users, ultimately sharpening their competitive edge. The platform also provides enterprise-grade security and infrastructure reliability for real-world deployments, backed by direct-to-expert support.

FPT AI Factory - Trusted regional AI infrastructure with expert-backed support
By combining its regional infrastructure with production-ready AI platforms, FPT AI Factory enables businesses across Southeast Asia and Japan to build and scale AI systems with greater performance, higher price-performance efficiency, and control in the inference era.
With NVIDIA HGX B300 GPU Cloud now officially available on FPT AI Factory, FPT is strengthening its position as enterprises move beyond model building toward large-scale AI deployment.