Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Horace He

Facebook AI

Flex Attention: A Programming Model for Generating Optimized Attention Kernels

Dec 07, 2024

Juechu Dong, Boyuan Feng, Driss Guessous, Yanbo Liang, Horace He

Abstract:Over the past 7 years, attention has become one of the most important primitives in deep learning. The primary approach to optimize attention is FlashAttention, which fuses the operation together, drastically improving both the runtime and the memory consumption. However, the importance of FlashAttention combined with its monolithic nature poses a problem for researchers aiming to try new attention variants -- a "software lottery". This problem is exacerbated by the difficulty of writing efficient fused attention kernels, resisting traditional compiler-based approaches. We introduce FlexAttention, a novel compiler-driven programming model that allows implementing the majority of attention variants in a few lines of idiomatic PyTorch code. We demonstrate that many existing attention variants (e.g. Alibi, Document Masking, PagedAttention, etc.) can be implemented via FlexAttention, and that we achieve competitive performance compared to these handwritten kernels. Finally, we demonstrate how FlexAttention allows for easy composition of attention variants, solving the combinatorial explosion of attention variants.

Via

Access Paper or Ask Questions

GPT-NeoX-20B: An Open-Source Autoregressive Language Model

Apr 14, 2022

Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang(+7 more)

Figure 1 for GPT-NeoX-20B: An Open-Source Autoregressive Language Model

Figure 2 for GPT-NeoX-20B: An Open-Source Autoregressive Language Model

Figure 3 for GPT-NeoX-20B: An Open-Source Autoregressive Language Model

Figure 4 for GPT-NeoX-20B: An Open-Source Autoregressive Language Model

Abstract:We introduce GPT-NeoX-20B, a 20 billion parameter autoregressive language model trained on the Pile, whose weights will be made freely and openly available to the public through a permissive license. It is, to the best of our knowledge, the largest dense autoregressive model that has publicly available weights at the time of submission. In this work, we describe \model{}'s architecture and training and evaluate its performance on a range of language-understanding, mathematics, and knowledge-based tasks. We find that GPT-NeoX-20B is a particularly powerful few-shot reasoner and gains far more in performance when evaluated five-shot than similarly sized GPT-3 and FairSeq models. We open-source the training and evaluation code, as well as the model weights, at https://github.com/EleutherAI/gpt-neox.

* To appear in the Proceedings of the ACL Workshop on Challenges & Perspectives in Creating Large Language Models

Via

Access Paper or Ask Questions

torch.fx: Practical Program Capture and Transformation for Deep Learning in Python

Dec 15, 2021

James K. Reed, Zachary DeVito, Horace He, Ansley Ussery, Jason Ansel

Figure 1 for torch.fx: Practical Program Capture and Transformation for Deep Learning in Python

Figure 2 for torch.fx: Practical Program Capture and Transformation for Deep Learning in Python

Figure 3 for torch.fx: Practical Program Capture and Transformation for Deep Learning in Python

Figure 4 for torch.fx: Practical Program Capture and Transformation for Deep Learning in Python

Abstract:Modern deep learning frameworks provide imperative, eager execution programming interfaces embedded in Python to provide a productive development experience. However, deep learning practitioners sometimes need to capture and transform program structure for performance optimization, visualization, analysis, and hardware integration. We study the different designs for program capture and transformation used in deep learning. By designing for typical deep learning use cases rather than long tail ones, it is possible to create a simpler framework for program capture and transformation. We apply this principle in torch.fx, a program capture and transformation library for PyTorch written entirely in Python and optimized for high developer productivity by ML practitioners. We present case studies showing how torch.fx enables workflows previously inaccessible in the PyTorch ecosystem.

* 14 pages, 8 figures, Submitted to MLSys 2022

Via

Access Paper or Ask Questions

Edge Proposal Sets for Link Prediction

Jun 30, 2021

Abhay Singh, Qian Huang, Sijia Linda Huang, Omkar Bhalerao, Horace He, Ser-Nam Lim, Austin R. Benson

Figure 1 for Edge Proposal Sets for Link Prediction

Figure 2 for Edge Proposal Sets for Link Prediction

Figure 3 for Edge Proposal Sets for Link Prediction

Figure 4 for Edge Proposal Sets for Link Prediction

Abstract:Graphs are a common model for complex relational data such as social networks and protein interactions, and such data can evolve over time (e.g., new friendships) and be noisy (e.g., unmeasured interactions). Link prediction aims to predict future edges or infer missing edges in the graph, and has diverse applications in recommender systems, experimental design, and complex systems. Even though link prediction algorithms strongly depend on the set of edges in the graph, existing approaches typically do not modify the graph topology to improve performance. Here, we demonstrate how simply adding a set of edges, which we call a \emph{proposal set}, to the graph as a pre-processing step can improve the performance of several link prediction algorithms. The underlying idea is that if the edges in the proposal set generally align with the structure of the graph, link prediction algorithms are further guided towards predicting the right edges; in other words, adding a proposal set of edges is a signal-boosting pre-processing step. We show how to use existing link prediction algorithms to generate effective proposal sets and evaluate this approach on various synthetic and empirical datasets. We find that proposal sets meaningfully improve the accuracy of link prediction algorithms based on both neighborhood heuristics and graph neural networks. Code is available at \url{https://github.com/CUAI/Edge-Proposal-Sets}.

Via

Access Paper or Ask Questions

Measuring Coding Challenge Competence With APPS

May 27, 2021

Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song(+1 more)

Figure 1 for Measuring Coding Challenge Competence With APPS

Figure 2 for Measuring Coding Challenge Competence With APPS

Figure 3 for Measuring Coding Challenge Competence With APPS

Figure 4 for Measuring Coding Challenge Competence With APPS

Abstract:While programming is one of the most broadly applicable skills in modern society, modern machine learning models still cannot code solutions to basic problems. Despite its importance, there has been surprisingly little work on evaluating code generation, and it can be difficult to accurately assess code generation performance rigorously. To meet this challenge, we introduce APPS, a benchmark for code generation. Unlike prior work in more restricted settings, our benchmark measures the ability of models to take an arbitrary natural language specification and generate satisfactory Python code. Similar to how companies assess candidate software developers, we then evaluate models by checking their generated code on test cases. Our benchmark includes 10,000 problems, which range from having simple one-line solutions to being substantial algorithmic challenges. We fine-tune large language models on both GitHub and our training set, and we find that the prevalence of syntax errors is decreasing exponentially as models improve. Recent models such as GPT-Neo can pass approximately 20% of the test cases of introductory problems, so we find that machine learning models are now beginning to learn how to code. As the social significance of automatic code generation increases over the coming years, our benchmark can provide an important measure for tracking advancements.

* Code and the APPS dataset is available at https://github.com/hendrycks/apps

Via

Access Paper or Ask Questions

The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Dec 31, 2020

Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima(+2 more)

Figure 1 for The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Figure 2 for The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Figure 3 for The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Figure 4 for The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Abstract:Recent work has demonstrated that increased training dataset diversity improves general cross-domain knowledge and downstream generalization capability for large-scale language models. With this in mind, we present \textit{the Pile}: an 825 GiB English text corpus targeted at training large-scale language models. The Pile is constructed from 22 diverse high-quality subsets -- both existing and newly constructed -- many of which derive from academic or professional sources. Our evaluation of the untuned performance of GPT-2 and GPT-3 on the Pile shows that these models struggle on many of its components, such as academic writing. Conversely, models trained on the Pile improve significantly over both Raw CC and CC-100 on all components of the Pile, while improving performance on downstream evaluations. Through an in-depth exploratory analysis, we document potentially concerning aspects of the data for prospective users. We make publicly available the code used in its construction.

Via

Access Paper or Ask Questions

Value Function Based Performance Optimization of Deep Learning Workloads

Nov 30, 2020

Benoit Steiner, Chris Cummins, Horace He, Hugh Leather

Figure 1 for Value Function Based Performance Optimization of Deep Learning Workloads

Figure 2 for Value Function Based Performance Optimization of Deep Learning Workloads

Figure 3 for Value Function Based Performance Optimization of Deep Learning Workloads

Abstract:As machine learning techniques become ubiquitous, the efficiency of neural network implementations is becoming correspondingly paramount. Frameworks, such as Halide and TVM, separate out the algorithmic representation of the network from the schedule that determines its implementation. Finding good schedules, however, remains extremely challenging. We model this scheduling problem as a sequence of optimization choices, and present a new technique to accurately predict the expected performance of a partial schedule. By leveraging these predictions we can make these optimization decisions greedily and rapidly identify an efficient schedule. This enables us to find schedules that improve the throughput of deep neural networks by 2.6x over Halide and 1.5x over TVM. Moreover, our technique is two to three orders of magnitude faster than that of these tools, and completes in seconds instead of hours.

Via

Access Paper or Ask Questions

Combining Label Propagation and Simple Models Out-performs Graph Neural Networks

Nov 02, 2020

Qian Huang, Horace He, Abhay Singh, Ser-Nam Lim, Austin R. Benson

Figure 1 for Combining Label Propagation and Simple Models Out-performs Graph Neural Networks

Figure 2 for Combining Label Propagation and Simple Models Out-performs Graph Neural Networks

Figure 3 for Combining Label Propagation and Simple Models Out-performs Graph Neural Networks

Figure 4 for Combining Label Propagation and Simple Models Out-performs Graph Neural Networks

Abstract:Graph Neural Networks (GNNs) are the predominant technique for learning over graphs. However, there is relatively little understanding of why GNNs are successful in practice and whether they are necessary for good performance. Here, we show that for many standard transductive node classification benchmarks, we can exceed or match the performance of state-of-the-art GNNs by combining shallow models that ignore the graph structure with two simple post-processing steps that exploit correlation in the label structure: (i) an "error correlation" that spreads residual errors in training data to correct errors in test data and (ii) a "prediction correlation" that smooths the predictions on the test data. We call this overall procedure Correct and Smooth (C&S), and the post-processing steps are implemented via simple modifications to standard label propagation techniques from early graph-based semi-supervised learning methods. Our approach exceeds or nearly matches the performance of state-of-the-art GNNs on a wide variety of benchmarks, with just a small fraction of the parameters and orders of magnitude faster runtime. For instance, we exceed the best known GNN performance on the OGB-Products dataset with 137 times fewer parameters and greater than 100 times less training time. The performance of our methods highlights how directly incorporating label information into the learning algorithm (as was done in traditional techniques) yields easy and substantial performance gains. We can also incorporate our techniques into big GNN models, providing modest gains. Our code for the OGB results is at https://github.com/Chillee/CorrectAndSmooth.

Via

Access Paper or Ask Questions

Set-Structured Latent Representations

Mar 09, 2020

Qian Huang, Horace He, Abhay Singh, Yan Zhang, Ser-Nam Lim, Austin Benson

Figure 1 for Set-Structured Latent Representations

Figure 2 for Set-Structured Latent Representations

Figure 3 for Set-Structured Latent Representations

Figure 4 for Set-Structured Latent Representations

Abstract:Unstructured data often has latent component structure, such as the objects in an image of a scene. In these situations, the relevant latent structure is an unordered collection or \emph{set}. However, learning such representations directly from data is difficult due to the discrete and unordered structure. Here, we develop a framework for differentiable learning of set-structured latent representations. We show how to use this framework to naturally decompose data such as images into sets of interpretable and meaningful components and demonstrate how existing techniques cannot properly disentangle relevant structure. We also show how to extend our methodology to downstream tasks such as set matching, which uses set-specific operations. Our code is available at https://github.com/CUVL/SSLR.

* Preprint, 17 pages

Via

Access Paper or Ask Questions

Enhancing Adversarial Example Transferability with an Intermediate Level Attack

Jul 23, 2019

Qian Huang, Isay Katsman, Horace He, Zeqi Gu, Serge Belongie, Ser-Nam Lim

Figure 1 for Enhancing Adversarial Example Transferability with an Intermediate Level Attack

Figure 2 for Enhancing Adversarial Example Transferability with an Intermediate Level Attack

Figure 3 for Enhancing Adversarial Example Transferability with an Intermediate Level Attack

Figure 4 for Enhancing Adversarial Example Transferability with an Intermediate Level Attack

Abstract:Neural networks are vulnerable to adversarial examples, malicious inputs crafted to fool trained models. Adversarial examples often exhibit black-box transfer, meaning that adversarial examples for one model can fool another model. However, adversarial examples are typically overfit to exploit the particular architecture and feature representation of a source model, resulting in sub-optimal black-box transfer attacks to other target models. We introduce the Intermediate Level Attack (ILA), which attempts to fine-tune an existing adversarial example for greater black-box transferability by increasing its perturbation on a pre-specified layer of the source model, improving upon state-of-the-art methods. We show that we can select a layer of the source model to perturb without any knowledge of the target models while achieving high transferability. Additionally, we provide some explanatory insights regarding our method and the effect of optimizing for adversarial examples in intermediate feature maps.

* ICCV 2019. arXiv admin note: text overlap with arXiv:1811.08458

Via

Access Paper or Ask Questions