Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Victor Cărbune

Can LLMs get help from other LLMs without revealing private information?

Apr 02, 2024

Florian Hartmann, Duc-Hieu Tran, Peter Kairouz, Victor Cărbune, Blaise Aguera y Arcas

Figure 1 for Can LLMs get help from other LLMs without revealing private information?

Figure 2 for Can LLMs get help from other LLMs without revealing private information?

Figure 3 for Can LLMs get help from other LLMs without revealing private information?

Figure 4 for Can LLMs get help from other LLMs without revealing private information?

Abstract:Cascades are a common type of machine learning systems in which a large, remote model can be queried if a local model is not able to accurately label a user's data by itself. Serving stacks for large language models (LLMs) increasingly use cascades due to their ability to preserve task performance while dramatically reducing inference costs. However, applying cascade systems in situations where the local model has access to sensitive data constitutes a significant privacy risk for users since such data could be forwarded to the remote model. In this work, we show the feasibility of applying cascade systems in such setups by equipping the local model with privacy-preserving techniques that reduce the risk of leaking private information when querying the remote model. To quantify information leakage in such setups, we introduce two privacy measures. We then propose a system that leverages the recently introduced social learning paradigm in which LLMs collaboratively learn from each other by exchanging natural language. Using this paradigm, we demonstrate on several datasets that our methods minimize the privacy loss while at the same time improving task performance compared to a non-cascade baseline.

Via

Access Paper or Ask Questions

ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Feb 19, 2024

Gilles Baechler, Srinivas Sunkara, Maria Wang, Fedir Zubach, Hassan Mansoor, Vincent Etter, Victor Cărbune, Jason Lin, Jindong Chen, Abhanshu Sharma

Figure 1 for ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Figure 2 for ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Figure 3 for ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Figure 4 for ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Abstract:Screen user interfaces (UIs) and infographics, sharing similar visual language and design principles, play important roles in human communication and human-machine interaction. We introduce ScreenAI, a vision-language model that specializes in UI and infographics understanding. Our model improves upon the PaLI architecture with the flexible patching strategy of pix2struct and is trained on a unique mixture of datasets. At the heart of this mixture is a novel screen annotation task in which the model has to identify the type and location of UI elements. We use these text annotations to describe screens to Large Language Models and automatically generate question-answering (QA), UI navigation, and summarization training datasets at scale. We run ablation studies to demonstrate the impact of these design choices. At only 5B parameters, ScreenAI achieves new state-of-the-artresults on UI- and infographics-based tasks (Multi-page DocVQA, WebSRC, MoTIF and Widget Captioning), and new best-in-class performance on others (Chart QA, DocVQA, and InfographicVQA) compared to models of similar size. Finally, we release three new datasets: one focused on the screen annotation task and two others focused on question answering.

* Revision notes: 1) In Appendix I, added dataset location for ScreenQA Short in Appendix I. 2) In Table 4, updated evaluation numbers for Screen Annotation and Complex Screen QA benchmarks as the datasets are updated. 3) Updated Figure 4 to reflect the changes in evaluation numbers described in 2). 4) Minor revisions in other places

Via

Access Paper or Ask Questions

LLMs cannot find reasoning errors, but can correct them!

Nov 14, 2023

Gladys Tyen, Hassan Mansoor, Peter Chen, Tony Mak, Victor Cărbune

Figure 1 for LLMs cannot find reasoning errors, but can correct them!

Figure 2 for LLMs cannot find reasoning errors, but can correct them!

Figure 3 for LLMs cannot find reasoning errors, but can correct them!

Figure 4 for LLMs cannot find reasoning errors, but can correct them!

Abstract:While self-correction has shown promise in improving LLM outputs in terms of style and quality (e.g. Chen et al., 2023; Madaan et al., 2023), recent attempts to self-correct logical or reasoning errors often cause correct answers to become incorrect, resulting in worse performances overall (Huang et al., 2023). In this paper, we break down the self-correction process into two core components: mistake finding and output correction. For mistake finding, we release BIG-Bench Mistake, a dataset of logical mistakes in Chain-of-Thought reasoning traces. We provide benchmark numbers for several state-of-the-art LLMs, and demonstrate that LLMs generally struggle with finding logical mistakes. For output correction, we propose a backtracking method which provides large improvements when given information on mistake location. We construe backtracking as a lightweight alternative to reinforcement learning methods, and show that it remains effective with a reward model at 60-70% accuracy.

Via

Access Paper or Ask Questions

Replacing Human Audio with Synthetic Audio for On-device Unspoken Punctuation Prediction

Oct 20, 2020

Daria Soboleva, Ondrej Skopek, Márius Šajgalík, Victor Cărbune, Felix Weissenberger, Julia Proskurnia, Bogdan Prisacari, Daniel Valcarce, Justin Lu, Rohit Prabhavalkar(+1 more)

Figure 1 for Replacing Human Audio with Synthetic Audio for On-device Unspoken Punctuation Prediction

Figure 2 for Replacing Human Audio with Synthetic Audio for On-device Unspoken Punctuation Prediction

Figure 3 for Replacing Human Audio with Synthetic Audio for On-device Unspoken Punctuation Prediction

Figure 4 for Replacing Human Audio with Synthetic Audio for On-device Unspoken Punctuation Prediction

Abstract:We present a novel multi-modal unspoken punctuation prediction system for the English language which combines acoustic and text features. We demonstrate for the first time, that by relying exclusively on synthetic data generated using a prosody-aware text-to-speech system, we can outperform a model trained with expensive human audio recordings on the unspoken punctuation prediction problem. Our model architecture is well suited for on-device use. This is achieved by leveraging hash-based embeddings of automatic speech recognition text output in conjunction with acoustic features as input to a quasi-recurrent neural network, keeping the model size small and latency low.

Via

Access Paper or Ask Questions