LPFormer: LiDAR Pose Estimation Transformer with Multi-Task Network

Ye, Dongqiangzi; Xie, Yufei; Chen, Weijia; Zhou, Zixiang; Ge, Lingting; Foroosh, Hassan

Computer Science > Computer Vision and Pattern Recognition

arXiv:2306.12525 (cs)

[Submitted on 21 Jun 2023 (v1), last revised 2 Mar 2024 (this version, v2)]

Title:LPFormer: LiDAR Pose Estimation Transformer with Multi-Task Network

Authors:Dongqiangzi Ye, Yufei Xie, Weijia Chen, Zixiang Zhou, Lingting Ge, Hassan Foroosh

View PDF HTML (experimental)

Abstract:Due to the difficulty of acquiring large-scale 3D human keypoint annotation, previous methods for 3D human pose estimation (HPE) have often relied on 2D image features and sequential 2D annotations. Furthermore, the training of these networks typically assumes the prediction of a human bounding box and the accurate alignment of 3D point clouds with 2D images, making direct application in real-world scenarios challenging. In this paper, we present the 1st framework for end-to-end 3D human pose estimation, named LPFormer, which uses only LiDAR as its input along with its corresponding 3D annotations. LPFormer consists of two stages: firstly, it identifies the human bounding box and extracts multi-level feature representations, and secondly, it utilizes a transformer-based network to predict human keypoints based on these features. Our method demonstrates that 3D HPE can be seamlessly integrated into a strong LiDAR perception network and benefit from the features extracted by the network. Experimental results on the Waymo Open Dataset demonstrate the state-of-the-art performance, and improvements even compared to previous multi-modal solutions.

Comments:	ICRA 2024. Top solution for the Waymo Open Dataset Challenges 2023 - Pose Estimation. CVPR 2023 Workshop on Autonomous Driving
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2306.12525 [cs.CV]
	(or arXiv:2306.12525v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2306.12525

Submission history

From: Zixiang Zhou [view email]
[v1] Wed, 21 Jun 2023 19:20:15 UTC (1,942 KB)
[v2] Sat, 2 Mar 2024 22:36:04 UTC (1,995 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:LPFormer: LiDAR Pose Estimation Transformer with Multi-Task Network

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:LPFormer: LiDAR Pose Estimation Transformer with Multi-Task Network

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators