In today's fast-evolving technology landscape, artificial intelligence (AI) is no longer just a passing trend but a strategic necessity for maintaining a competitive edge. Businesses across sectors are under increasing pressure to harness AI to drive efficiency, innovation, and smarter decision-making.

Against this backdrop, the tech industry is now turning its attention to Large Vision Models (LVMs) as the next leap forward. These models are designed to transform how visual data is recognized and analyzed, unlocking new possibilities in image-based insights. In particular, LVMs are poised to reshape critical fields such as healthcare and automotive, which will be the focus of this article.

What is a Large Vision Model?

A large vision model (LVM) is an advanced artificial intelligence model built to handle complex visual tasks, such as image recognition, object detection, segmentation, and image generation. It is designed to understand and create visual content from images and videos, in a similar way that language models work with text.

LVMs are typically based on Vision Transformer architectures and are trained on massive datasets of images and videos. Through this large-scale training, they learn to recognize patterns, classify objects, and generate new visual content. Vision Transformers enable LVMs to generalize effectively from vast training data, which in turn gives them strong few-shot and zero-shot performance across a wide range of downstream visual tasks.

Healthcare applications of Large Vision Models

Healthcare providers often struggle to detect subtle abnormalities because human anatomical structures have highly complex shapes and textures. Large Vision Models, particularly Vision Transformers (ViT), are transforming medical imaging workflows through several key applications:

  • Image reconstruction and synthesis: Medical imaging data such as MRI and CT scans are frequently stored in unstructured formats, making it difficult to reconstruct clear, detailed images. Traditional reconstruction methods rely on complex algorithms and can be time-consuming. Vision Transformers accelerate and enhance image reconstruction by splitting images into small patches (tokens) and selectively focusing on the most relevant regions and features. This token-based attention helps preserve critical details while producing high-quality reconstructions within seconds.
  • Image segmentation: Research by Chen et al. (2021) introduced TransUNet, one of the earliest models that combines Vision Transformers (ViT) with the UNet architecture for medical image segmentation. UNet excels at precise object segmentation and preserving fine details but has limitations in handling sequence-to-sequence relationships and complex feature extraction. ViT is strong in modeling sequence-to-sequence features yet lacks fine-grained feature localization. By merging these two architectures, TransUNet delivers effective multiorgan segmentation, which is essential for analyzing complex structures in MRI and CT images.
  • Surgical scene reconstruction: In surgical environments, ViT-based stereo transformers are used to reconstruct dynamic surgical scenes, supporting surgical education, robot-assisted interventions, and context-aware decision-making. A study by Wang et al. (2021) employed a Swin Transformer model to accurately reconstruct sinograms from CT scans. This approach produces high-quality images, enables lower radiation doses, and supports earlier cancer detection.
  • Medical report generation: Vision Transformer–based frameworks can assist clinicians in generating radiology reports, surgical instructions, and other clinical documentation by leveraging large volumes of data stored in health information systems. This technology helps address issues such as biased medical data and long, inconsistent narrative reports. You et al. (2021) proposed the AlignTransformer framework, which produces long, descriptive, and coherent paragraphs based on medical image analysis. AlignTransformer operates in two stages: first, it aligns medical tags with corresponding images to extract relevant features; second, it uses these features, together with training data for each medical tag, to generate comprehensive medical reports.

Recognizing the potential of computer vision, FPT has strategically partnered with Landing AI, a leading computer vision and AI software company based in the United States, to embed advanced computer vision capabilities into its ecosystem.

At the “Visionary Integration: Showcasing the Future of Computer Vision” workshop in March 2024, FPT’s Chief AI Transformation Officer, Dr. Nguyen Xuan Phong, emphasized the importance of computer vision integration across FPT’s operations. He highlighted that Landing AI provides some of the world’s most advanced computer vision technology based on transformer architectures, with broad compatibility, and stressed the need to harness this technology to maximize value for FPT.

FPT has already piloted Landing AI’s solutions in the healthcare domain with encouraging outcomes. In one deployment, Landing AI supported the diagnosis of 30 dermatological diseases using a model trained on 8,844 images, achieving an accuracy rate of 93%. According to FPT, Landing AI’s current imaging technology has enhanced diagnostic performance by a factor of 10 compared with previous imaging solutions.

Large Vision Models in the Automotive Industry

In the automotive industry, advanced vision models play a critical role in enhancing driver assistance systems and delivering safer driving experiences. A central task in these systems is the detection and classification of vehicles, which is essential for Advanced Driver Assistance Systems (ADAS) and Intelligent Transportation Systems (ITS) that aim to reduce road accidents and save lives.

According to the World Health Organization (WHO), approximately 1.19 million lives are lost every year due to road traffic crashes, with millions more suffering non-fatal injuries and long-term disabilities.

To address this pressing safety challenge, Taki and Zemmouri (2023) proposed an innovative vehicle image classification method based on a Vision Transformer model. They leveraged a pre-trained Vision Transformer, one of the latest advances in computer vision, to tackle the vehicle classification problem under challenging conditions where traditional object detection models often struggle, such as low-quality images, nighttime scenes, and poor illumination.

The researchers used the ImageNet-21k dataset, working with 4,800 tiny, low-resolution vehicle images categorized into six classes: Bike, Car, Juggernaut, Minibus, Pickup, and Truck. Their approach achieved a promising accuracy of 99.3% using the Vision Transformer, surpassing previous methods on vehicle classification tasks.

In a separate real-world application, FPT utilized Landing AI’s solutions to help a major car interior supplier control the quality of car door assembly. The solution was implemented within just one month, covering data collection, labeling, model training, and deployment. Thanks to this solution, the company achieved 99.7% accuracy in defect detection and significantly optimized quality control time, reducing inspection duration from 3 minutes to only 2 seconds.

Large Vision Models: Enlightening the Tech Landscape

The integration of large vision models with existing Large Language Models is becoming increasingly crucial as the technology industry evolves. This powerful combination enables comprehensive AI systems that can seamlessly navigate and understand both textual and visual information. It also supports more natural human interaction with AI systems, whether through text or voice.

With over a decade of investment in AI, FPT has achieved significant milestones. The company has built an AI ecosystem with more than 20 products and solutions, serving over 20 million users across 15 countries. Building on this foundation, FPT continues to harness the power of computer vision across industries for long-term, sustainable impact.