Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Artem Chumachenko

Distributed Inference and Fine-tuning of Large Language Models Over The Internet

Dec 13, 2023

Alexander Borzunov, Max Ryabinin, Artem Chumachenko, Dmitry Baranchuk, Tim Dettmers, Younes Belkada, Pavel Samygin, Colin Raffel

Figure 1 for Distributed Inference and Fine-tuning of Large Language Models Over The Internet

Figure 2 for Distributed Inference and Fine-tuning of Large Language Models Over The Internet

Figure 3 for Distributed Inference and Fine-tuning of Large Language Models Over The Internet

Figure 4 for Distributed Inference and Fine-tuning of Large Language Models Over The Internet

Abstract:Large language models (LLMs) are useful in many NLP tasks and become more capable with size, with the best open-source models having over 50 billion parameters. However, using these 50B+ models requires high-end hardware, making them inaccessible to most researchers. In this work, we investigate methods for cost-efficient inference and fine-tuning of LLMs, comparing local and distributed strategies. We observe that a large enough model (50B+) can run efficiently even on geodistributed devices in a consumer-grade network. This could allow running LLM efficiently by pooling together idle compute resources of multiple research groups and volunteers. We address two open problems: (1) how to perform inference and fine-tuning reliably if any device can disconnect abruptly and (2) how to partition LLMs between devices with uneven hardware, joining and leaving at will. In order to do that, we develop special fault-tolerant inference algorithms and load-balancing protocols that automatically assign devices to maximize the total system throughput. We showcase these algorithms in Petals - a decentralized system that runs Llama 2 (70B) and BLOOM (176B) over the Internet up to 10x faster than offloading for interactive generation. We evaluate the performance of our system in simulated conditions and a real-world setup spanning two continents.

* Accepted to Conference on Neural Information Processing Systems (NeurIPS) 2023. 20 pages, 3 figures

Via

Access Paper or Ask Questions

Petals: Collaborative Inference and Fine-tuning of Large Models

Sep 02, 2022

Alexander Borzunov, Dmitry Baranchuk, Tim Dettmers, Max Ryabinin, Younes Belkada, Artem Chumachenko, Pavel Samygin, Colin Raffel

Figure 1 for Petals: Collaborative Inference and Fine-tuning of Large Models

Figure 2 for Petals: Collaborative Inference and Fine-tuning of Large Models

Figure 3 for Petals: Collaborative Inference and Fine-tuning of Large Models

Figure 4 for Petals: Collaborative Inference and Fine-tuning of Large Models

Abstract:Many NLP tasks benefit from using large language models (LLMs) that often have more than 100 billion parameters. With the release of BLOOM-176B and OPT-175B, everyone can download pretrained models of this scale. Still, using these models requires high-end hardware unavailable to many researchers. In some cases, LLMs can be used more affordably via RAM offloading or hosted APIs. However, these techniques have innate limitations: offloading is too slow for interactive inference, while APIs are not flexible enough for research. In this work, we propose Petals $-$ a system for inference and fine-tuning of large models collaboratively by joining the resources of multiple parties trusted to process client's data. We demonstrate that this strategy significantly outperforms offloading for very large models, running inference of BLOOM-176B on consumer GPUs with $\approx$ 1 step per second. Unlike most inference APIs, Petals also natively exposes the hidden states of served models, allowing its users to train and share custom model extensions based on efficient fine-tuning methods.

* 10 pages, 4 figures. Link: https://petals.ml

Via

Access Paper or Ask Questions

Weight Squeezing: Reparameterization for Compression and Fast Inference

Oct 14, 2020

Artem Chumachenko, Daniil Gavrilov, Pavel Kalaidin

Figure 1 for Weight Squeezing: Reparameterization for Compression and Fast Inference

Figure 2 for Weight Squeezing: Reparameterization for Compression and Fast Inference

Figure 3 for Weight Squeezing: Reparameterization for Compression and Fast Inference

Figure 4 for Weight Squeezing: Reparameterization for Compression and Fast Inference

Abstract:In this work, we present a novel approach for simultaneous knowledge transfer and model compression called Weight Squeezing. With this method, we perform knowledge transfer from a pre-trained teacher model by learning the mapping from its weights to smaller student model weights, without significant loss of model accuracy. We applied Weight Squeezing combined with Knowledge Distillation to a pre-trained text classification model, and compared it to various knowledge transfer and model compression methods on several downstream text classification tasks. We observed that our approach produces better results than Knowledge Distillation methods without any loss in inference speed. We also compared Weight Squeezing with Low Rank Factorization methods and observed that our method is significantly faster at inference while being competitive in terms of accuracy.

Via

Access Paper or Ask Questions