Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:PP-DocLayout: A Unified Document Layout Detection Model to Accelerate Large-Scale Data Construction

Mar 21, 2025

Ting Sun, Cheng Cui, Yuning Du, Yi Liu

Figure 1 for PP-DocLayout: A Unified Document Layout Detection Model to Accelerate Large-Scale Data Construction

Figure 2 for PP-DocLayout: A Unified Document Layout Detection Model to Accelerate Large-Scale Data Construction

Figure 3 for PP-DocLayout: A Unified Document Layout Detection Model to Accelerate Large-Scale Data Construction

Figure 4 for PP-DocLayout: A Unified Document Layout Detection Model to Accelerate Large-Scale Data Construction

Share this with someone who'll enjoy it:

Abstract:Document layout analysis is a critical preprocessing step in document intelligence, enabling the detection and localization of structural elements such as titles, text blocks, tables, and formulas. Despite its importance, existing layout detection models face significant challenges in generalizing across diverse document types, handling complex layouts, and achieving real-time performance for large-scale data processing. To address these limitations, we present PP-DocLayout, which achieves high precision and efficiency in recognizing 23 types of layout regions across diverse document formats. To meet different needs, we offer three models of varying scales. PP-DocLayout-L is a high-precision model based on the RT-DETR-L detector, achieving 90.4% mAP@0.5 and an end-to-end inference time of 13.4 ms per page on a T4 GPU. PP-DocLayout-M is a balanced model, offering 75.2% mAP@0.5 with an inference time of 12.7 ms per page on a T4 GPU. PP-DocLayout-S is a high-efficiency model designed for resource-constrained environments and real-time applications, with an inference time of 8.1 ms per page on a T4 GPU and 14.5 ms on a CPU. This work not only advances the state of the art in document layout analysis but also provides a robust solution for constructing high-quality training data, enabling advancements in document intelligence and multimodal AI systems. Code and models are available at https://github.com/PaddlePaddle/PaddleX .

* Github Repo: https://github.com/PaddlePaddle/PaddleX

View paper on

Share this with someone who'll enjoy it:

Title:PP-DocLayout: A Unified Document Layout Detection Model to Accelerate Large-Scale Data Construction

Paper and Code