MARATTO

article

Point-Net Vision Transformer: An Extendable 3D Object Reconstruction Approach

Abstract

3D object reconstruction builds 3D objects from either a single image or many images of the object from different perspectives and can be done manually or digitally. Digital approaches include Pixel-Aligned Implicit Function (PIFu) and Mesh Generation Network (MGN) variants. Its digital approaches have many applications from effective treatment in medicine to efficient object building in engineering. Traditional 3D reconstruction relies on costly depth sensors, making it challenging and slow. However, implementing this task digitally can achieve similar results faster. This paper discusses an object reconstruction approach based on merging the PointNet and the Vision Transformer that reconstructs 3D point clouds of objects from different domains through single or multiple images of the objects without user intervention. The proposed model was trained on both the BUFF and the Pix3D datasets, tackling different domains for a generalizable approach over 100 epochs, having its best performance at 35 epochs, achieving a Chamfer distance score of 0.0324cm, making this model a state-of-the-art on both datasets, when compared to related works.

Research topics

  • Advanced Vision and Imaging
  • Robotics and Sensor-Based Localization
  • Optical measurement and interference techniques

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/icca66035.2025.11430993

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.