Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Siyang Qin

Unified Autoregressive Visual Generation and Understanding with Continuous Tokens

Mar 17, 2025

Lijie Fan, Luming Tang, Siyang Qin, Tianhong Li, Xuan Yang, Siyuan Qiao, Andreas Steiner, Chen Sun, Yuanzhen Li, Tao Zhu(+4 more)

Abstract:We present UniFluid, a unified autoregressive framework for joint visual generation and understanding leveraging continuous visual tokens. Our unified autoregressive architecture processes multimodal image and text inputs, generating discrete tokens for text and continuous tokens for image. We find though there is an inherent trade-off between the image generation and understanding task, a carefully tuned training recipe enables them to improve each other. By selecting an appropriate loss balance weight, the unified model achieves results comparable to or exceeding those of single-task baselines on both tasks. Furthermore, we demonstrate that employing stronger pre-trained LLMs and random-order generation during training is important to achieve high-fidelity image generation within this unified framework. Built upon the Gemma model series, UniFluid exhibits competitive performance across both image generation and understanding, demonstrating strong transferability to various downstream tasks, including image editing for generation, as well as visual captioning and question answering for understanding.

* Tech report

Via

Access Paper or Ask Questions

PaliGemma 2: A Family of Versatile VLMs for Transfer

Dec 04, 2024

Andreas Steiner, André Susano Pinto, Michael Tschannen, Daniel Keysers, Xiao Wang, Yonatan Bitton, Alexey Gritsenko, Matthias Minderer, Anthony Sherbondy, Shangbang Long(+8 more)

Figure 1 for PaliGemma 2: A Family of Versatile VLMs for Transfer

Figure 2 for PaliGemma 2: A Family of Versatile VLMs for Transfer

Figure 3 for PaliGemma 2: A Family of Versatile VLMs for Transfer

Figure 4 for PaliGemma 2: A Family of Versatile VLMs for Transfer

Abstract:PaliGemma 2 is an upgrade of the PaliGemma open Vision-Language Model (VLM) based on the Gemma 2 family of language models. We combine the SigLIP-So400m vision encoder that was also used by PaliGemma with the whole range of Gemma 2 models, from the 2B one all the way up to the 27B model. We train these models at three resolutions (224px, 448px, and 896px) in multiple stages to equip them with broad knowledge for transfer via fine-tuning. The resulting family of base models covering different model sizes and resolutions allows us to investigate factors impacting transfer performance (such as learning rate) and to analyze the interplay between the type of task, model size, and resolution. We further increase the number and breadth of transfer tasks beyond the scope of PaliGemma including different OCR-related tasks such as table structure recognition, molecular structure recognition, music score recognition, as well as long fine-grained captioning and radiography report generation, on which PaliGemma 2 obtains state-of-the-art results.

Via

Access Paper or Ask Questions

Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens

Oct 17, 2024

Lijie Fan, Tianhong Li, Siyang Qin, Yuanzhen Li, Chen Sun, Michael Rubinstein, Deqing Sun, Kaiming He, Yonglong Tian

Figure 1 for Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens

Figure 2 for Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens

Figure 3 for Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens

Figure 4 for Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens

Abstract:Scaling up autoregressive models in vision has not proven as beneficial as in large language models. In this work, we investigate this scaling problem in the context of text-to-image generation, focusing on two critical factors: whether models use discrete or continuous tokens, and whether tokens are generated in a random or fixed raster order using BERT- or GPT-like transformer architectures. Our empirical results show that, while all models scale effectively in terms of validation loss, their evaluation performance -- measured by FID, GenEval score, and visual quality -- follows different trends. Models based on continuous tokens achieve significantly better visual quality than those using discrete tokens. Furthermore, the generation order and attention mechanisms significantly affect the GenEval score: random-order models achieve notably better GenEval scores compared to raster-order models. Inspired by these findings, we train Fluid, a random-order autoregressive model on continuous tokens. Fluid 10.5B model achieves a new state-of-the-art zero-shot FID of 6.16 on MS-COCO 30K, and 0.69 overall score on the GenEval benchmark. We hope our findings and results will encourage future efforts to further bridge the scaling gap between vision and language models.

* Tech report

Via

Access Paper or Ask Questions

Hierarchical Text Spotter for Joint Text Spotting and Layout Analysis

Oct 25, 2023

Shangbang Long, Siyang Qin, Yasuhisa Fujii, Alessandro Bissacco, Michalis Raptis

Figure 1 for Hierarchical Text Spotter for Joint Text Spotting and Layout Analysis

Figure 2 for Hierarchical Text Spotter for Joint Text Spotting and Layout Analysis

Figure 3 for Hierarchical Text Spotter for Joint Text Spotting and Layout Analysis

Figure 4 for Hierarchical Text Spotter for Joint Text Spotting and Layout Analysis

Abstract:We propose Hierarchical Text Spotter (HTS), a novel method for the joint task of word-level text spotting and geometric layout analysis. HTS can recognize text in an image and identify its 4-level hierarchical structure: characters, words, lines, and paragraphs. The proposed HTS is characterized by two novel components: (1) a Unified-Detector-Polygon (UDP) that produces Bezier Curve polygons of text lines and an affinity matrix for paragraph grouping between detected lines; (2) a Line-to-Character-to-Word (L2C2W) recognizer that splits lines into characters and further merges them back into words. HTS achieves state-of-the-art results on multiple word-level text spotting benchmark datasets as well as geometric layout analysis tasks.

* Accepted to WACV 2024

Via

Access Paper or Ask Questions

ICDAR 2023 Competition on Hierarchical Text Detection and Recognition

May 16, 2023

Shangbang Long, Siyang Qin, Dmitry Panteleev, Alessandro Bissacco, Yasuhisa Fujii, Michalis Raptis

Figure 1 for ICDAR 2023 Competition on Hierarchical Text Detection and Recognition

Figure 2 for ICDAR 2023 Competition on Hierarchical Text Detection and Recognition

Figure 3 for ICDAR 2023 Competition on Hierarchical Text Detection and Recognition

Figure 4 for ICDAR 2023 Competition on Hierarchical Text Detection and Recognition

Abstract:We organize a competition on hierarchical text detection and recognition. The competition is aimed to promote research into deep learning models and systems that can jointly perform text detection and recognition and geometric layout analysis. We present details of the proposed competition organization, including tasks, datasets, evaluations, and schedule. During the competition period (from January 2nd 2023 to April 1st 2023), at least 50 submissions from more than 20 teams were made in the 2 proposed tasks. Considering the number of teams and submissions, we conclude that the HierText competition has been successfully held. In this report, we will also present the competition results and insights from them.

* ICDAR 2023 competition report by organizers (accepted and to be published officially later)

Via

Access Paper or Ask Questions

FormNetV2: Multimodal Graph Contrastive Learning for Form Document Information Extraction

May 04, 2023

Chen-Yu Lee, Chun-Liang Li, Hao Zhang, Timothy Dozat, Vincent Perot, Guolong Su, Xiang Zhang, Kihyuk Sohn, Nikolai Glushnev, Renshen Wang(+6 more)

Figure 1 for FormNetV2: Multimodal Graph Contrastive Learning for Form Document Information Extraction

Figure 2 for FormNetV2: Multimodal Graph Contrastive Learning for Form Document Information Extraction

Figure 3 for FormNetV2: Multimodal Graph Contrastive Learning for Form Document Information Extraction

Figure 4 for FormNetV2: Multimodal Graph Contrastive Learning for Form Document Information Extraction

Abstract:The recent advent of self-supervised pre-training techniques has led to a surge in the use of multimodal learning in form document understanding. However, existing approaches that extend the mask language modeling to other modalities require careful multi-task tuning, complex reconstruction target designs, or additional pre-training data. In FormNetV2, we introduce a centralized multimodal graph contrastive learning strategy to unify self-supervised pre-training for all modalities in one loss. The graph contrastive objective maximizes the agreement of multimodal representations, providing a natural interplay for all modalities without special customization. In addition, we extract image features within the bounding box that joins a pair of tokens connected by a graph edge, capturing more targeted visual cues without loading a sophisticated and separately pre-trained image embedder. FormNetV2 establishes new state-of-the-art performance on FUNSD, CORD, SROIE and Payment benchmarks with a more compact model size.

* Accepted to ACL 2023

Via

Access Paper or Ask Questions

Towards End-to-End Unified Scene Text Detection and Layout Analysis

Mar 28, 2022

Shangbang Long, Siyang Qin, Dmitry Panteleev, Alessandro Bissacco, Yasuhisa Fujii, Michalis Raptis

Figure 1 for Towards End-to-End Unified Scene Text Detection and Layout Analysis

Figure 2 for Towards End-to-End Unified Scene Text Detection and Layout Analysis

Figure 3 for Towards End-to-End Unified Scene Text Detection and Layout Analysis

Figure 4 for Towards End-to-End Unified Scene Text Detection and Layout Analysis

Abstract:Scene text detection and document layout analysis have long been treated as two separate tasks in different image domains. In this paper, we bring them together and introduce the task of unified scene text detection and layout analysis. The first hierarchical scene text dataset is introduced to enable this novel research task. We also propose a novel method that is able to simultaneously detect scene text and form text clusters in a unified way. Comprehensive experiments show that our unified model achieves better performance than multiple well-designed baseline methods. Additionally, this model achieves state-of-the-art results on multiple scene text detection datasets without the need of complex post-processing. Dataset and code: https://github.com/google-research-datasets/hiertext.

* To appear at CVPR 2022

Via

Access Paper or Ask Questions

ROPE: Reading Order Equivariant Positional Encoding for Graph-based Document Information Extraction

Jun 21, 2021

Chen-Yu Lee, Chun-Liang Li, Chu Wang, Renshen Wang, Yasuhisa Fujii, Siyang Qin, Ashok Popat, Tomas Pfister

Figure 1 for ROPE: Reading Order Equivariant Positional Encoding for Graph-based Document Information Extraction

Figure 2 for ROPE: Reading Order Equivariant Positional Encoding for Graph-based Document Information Extraction

Figure 3 for ROPE: Reading Order Equivariant Positional Encoding for Graph-based Document Information Extraction

Figure 4 for ROPE: Reading Order Equivariant Positional Encoding for Graph-based Document Information Extraction

Abstract:Natural reading orders of words are crucial for information extraction from form-like documents. Despite recent advances in Graph Convolutional Networks (GCNs) on modeling spatial layout patterns of documents, they have limited ability to capture reading orders of given word-level node representations in a graph. We propose Reading Order Equivariant Positional Encoding (ROPE), a new positional encoding technique designed to apprehend the sequential presentation of words in documents. ROPE generates unique reading order codes for neighboring words relative to the target word given a word-level graph connectivity. We study two fundamental document entity extraction tasks including word labeling and word grouping on the public FUNSD dataset and a large-scale payment dataset. We show that ROPE consistently improves existing GCNs with a margin up to 8.4% F1-score.

* Accepted to ACL-IJCNLP 2021 (Oral)

Via

Access Paper or Ask Questions

Rethinking Text Line Recognition Models

Apr 21, 2021

Daniel Hernandez Diaz, Siyang Qin, Reeve Ingle, Yasuhisa Fujii, Alessandro Bissacco

Figure 1 for Rethinking Text Line Recognition Models

Figure 2 for Rethinking Text Line Recognition Models

Figure 3 for Rethinking Text Line Recognition Models

Figure 4 for Rethinking Text Line Recognition Models

Abstract:In this paper, we study the problem of text line recognition. Unlike most approaches targeting specific domains such as scene-text or handwritten documents, we investigate the general problem of developing a universal architecture that can extract text from any image, regardless of source or input modality. We consider two decoder families (Connectionist Temporal Classification and Transformer) and three encoder modules (Bidirectional LSTMs, Self-Attention, and GRCLs), and conduct extensive experiments to compare their accuracy and performance on widely used public datasets of scene and handwritten text. We find that a combination that so far has received little attention in the literature, namely a Self-Attention encoder coupled with the CTC decoder, when compounded with an external language model and trained on both public and internal data, outperforms all the others in accuracy and computational complexity. Unlike the more common Transformer-based models, this architecture can handle inputs of arbitrary length, a requirement for universal line recognition. Using an internal dataset collected from multiple sources, we also expose the limitations of current public datasets in evaluating the accuracy of line recognizers, as the relatively narrow image width and sequence length distributions do not allow to observe the quality degradation of the Transformer approach when applied to the transcription of long lines.

* 11 pages, 6 figures

Via

Access Paper or Ask Questions

Towards Unconstrained End-to-End Text Spotting

Aug 24, 2019

Siyang Qin, Alessandro Bissacco, Michalis Raptis, Yasuhisa Fujii, Ying Xiao

Figure 1 for Towards Unconstrained End-to-End Text Spotting

Figure 2 for Towards Unconstrained End-to-End Text Spotting

Figure 3 for Towards Unconstrained End-to-End Text Spotting

Figure 4 for Towards Unconstrained End-to-End Text Spotting

Abstract:We propose an end-to-end trainable network that can simultaneously detect and recognize text of arbitrary shape, making substantial progress on the open problem of reading scene text of irregular shape. We formulate arbitrary shape text detection as an instance segmentation problem; an attention model is then used to decode the textual content of each irregularly shaped text region without rectification. To extract useful irregularly shaped text instance features from image scale features, we propose a simple yet effective RoI masking step. Additionally, we show that predictions from an existing multi-step OCR engine can be leveraged as partially labeled training data, which leads to significant improvements in both the detection and recognition accuracy of our model. Our method surpasses the state-of-the-art for end-to-end recognition tasks on the ICDAR15 (straight) benchmark by 4.6%, and on the Total-Text (curved) benchmark by more than 16%.

* Accepted to ICCV 2019 as oral presentation

Via

Access Paper or Ask Questions