Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Yonatan Geifman

Llama-Nemotron: Efficient Reasoning Models

May 02, 2025

Akhiad Bercovich, Itay Levy, Izik Golan, Mohammad Dabbah, Ran El-Yaniv, Omri Puny, Ido Galil, Zach Moshe, Tomer Ronen, Najeeb Nabwani(+121 more)

Abstract:We introduce the Llama-Nemotron series of models, an open family of heterogeneous reasoning models that deliver exceptional reasoning capabilities, inference efficiency, and an open license for enterprise use. The family comes in three sizes -- Nano (8B), Super (49B), and Ultra (253B) -- and performs competitively with state-of-the-art reasoning models such as DeepSeek-R1 while offering superior inference throughput and memory efficiency. In this report, we discuss the training procedure for these models, which entails using neural architecture search from Llama 3 models for accelerated inference, knowledge distillation, and continued pretraining, followed by a reasoning-focused post-training stage consisting of two main parts: supervised fine-tuning and large scale reinforcement learning. Llama-Nemotron models are the first open-source models to support a dynamic reasoning toggle, allowing users to switch between standard chat and reasoning modes during inference. To further support open research and facilitate model development, we provide the following resources: 1. We release the Llama-Nemotron reasoning models -- LN-Nano, LN-Super, and LN-Ultra -- under the commercially permissive NVIDIA Open Model License Agreement. 2. We release the complete post-training dataset: Llama-Nemotron-Post-Training-Dataset. 3. We also release our training codebases: NeMo, NeMo-Aligner, and Megatron-LM.

Via

Access Paper or Ask Questions

FFN Fusion: Rethinking Sequential Computation in Large Language Models

Mar 24, 2025

Akhiad Bercovich, Mohammad Dabbah, Omri Puny, Ido Galil, Amnon Geifman, Yonatan Geifman, Izhak Golan, Ehud Karpas, Itay Levy, Zach Moshe(+8 more)

Abstract:We introduce FFN Fusion, an architectural optimization technique that reduces sequential computation in large language models by identifying and exploiting natural opportunities for parallelization. Our key insight is that sequences of Feed-Forward Network (FFN) layers, particularly those remaining after the removal of specific attention layers, can often be parallelized with minimal accuracy impact. We develop a principled methodology for identifying and fusing such sequences, transforming them into parallel operations that significantly reduce inference latency while preserving model behavior. Applying these techniques to Llama-3.1-405B-Instruct, we create Llama-Nemotron-Ultra-253B-Base (Ultra-253B-Base), an efficient and soon-to-be publicly available model that achieves a 1.71X speedup in inference latency and 35X lower per-token cost while maintaining strong performance across benchmarks. Through extensive experiments on models from 49B to 253B parameters, we demonstrate that FFN Fusion becomes increasingly effective at larger scales and can complement existing optimization techniques like quantization and pruning. Most intriguingly, we find that even full transformer blocks containing both attention and FFN layers can sometimes be parallelized, suggesting new directions for neural architecture design.

Via

Access Paper or Ask Questions

Puzzle: Distillation-Based NAS for Inference-Optimized LLMs

Dec 03, 2024

Akhiad Bercovich, Tomer Ronen, Talor Abramovich, Nir Ailon, Nave Assaf, Mohammad Dabbah, Ido Galil, Amnon Geifman, Yonatan Geifman, Izhak Golan(+16 more)

Figure 1 for Puzzle: Distillation-Based NAS for Inference-Optimized LLMs

Figure 2 for Puzzle: Distillation-Based NAS for Inference-Optimized LLMs

Figure 3 for Puzzle: Distillation-Based NAS for Inference-Optimized LLMs

Figure 4 for Puzzle: Distillation-Based NAS for Inference-Optimized LLMs

Abstract:Large language models (LLMs) have demonstrated remarkable capabilities, but their adoption is limited by high computational costs during inference. While increasing parameter counts enhances accuracy, it also widens the gap between state-of-the-art capabilities and practical deployability. We present Puzzle, a framework to accelerate LLM inference on specific hardware while preserving their capabilities. Through an innovative application of neural architecture search (NAS) at an unprecedented scale, Puzzle systematically optimizes models with tens of billions of parameters under hardware constraints. Our approach utilizes blockwise local knowledge distillation (BLD) for parallel architecture exploration and employs mixed-integer programming for precise constraint optimization. We demonstrate the real-world impact of our framework through Llama-3.1-Nemotron-51B-Instruct (Nemotron-51B), a publicly available model derived from Llama-3.1-70B-Instruct. Nemotron-51B achieves a 2.17x inference throughput speedup, fitting on a single NVIDIA H100 GPU while preserving 98.4% of the original model's capabilities. Nemotron-51B currently stands as the most accurate language model capable of inference on a single GPU with large batch sizes. Remarkably, this transformation required just 45B training tokens, compared to over 15T tokens used for the 70B model it was derived from. This establishes a new paradigm where powerful models can be optimized for efficient deployment with only negligible compromise of their capabilities, demonstrating that inference performance, not parameter count alone, should guide model selection. With the release of Nemotron-51B and the presentation of the Puzzle framework, we provide practitioners immediate access to state-of-the-art language modeling capabilities at significantly reduced computational costs.

Via

Access Paper or Ask Questions

SelectiveNet: A Deep Neural Network with an Integrated Reject Option

Jan 26, 2019

Yonatan Geifman, Ran El-Yaniv

Figure 1 for SelectiveNet: A Deep Neural Network with an Integrated Reject Option

Figure 2 for SelectiveNet: A Deep Neural Network with an Integrated Reject Option

Figure 3 for SelectiveNet: A Deep Neural Network with an Integrated Reject Option

Figure 4 for SelectiveNet: A Deep Neural Network with an Integrated Reject Option

Abstract:We consider the problem of selective prediction (also known as reject option) in deep neural networks, and introduce SelectiveNet, a deep neural architecture with an integrated reject option. Existing rejection mechanisms are based mostly on a threshold over the prediction confidence of a pre-trained network. In contrast, SelectiveNet is trained to optimize both classification (or regression) and rejection simultaneously, end-to-end. The result is a deep neural network that is optimized over the covered domain. In our experiments, we show a consistently improved risk-coverage trade-off over several well-known classification and regression datasets, thus reaching new state-of-the-art results for deep selective classification.

Via

Access Paper or Ask Questions

Deep Active Learning with a Neural Architecture Search

Nov 19, 2018

Yonatan Geifman, Ran El-Yaniv

Figure 1 for Deep Active Learning with a Neural Architecture Search

Figure 2 for Deep Active Learning with a Neural Architecture Search

Figure 3 for Deep Active Learning with a Neural Architecture Search

Figure 4 for Deep Active Learning with a Neural Architecture Search

Abstract:We consider active learning of deep neural networks. Most active learning works in this context have focused on studying effective querying mechanisms and assumed that an appropriate network architecture is a priori known for the problem at hand. We challenge this assumption and propose a novel active strategy whereby the learning algorithm searches for effective architectures on the fly, while actively learning. We apply our strategy using three known querying techniques (softmax response, MC-dropout, and coresets) and show that the proposed approach overwhelmingly outperforms active learning using fixed architectures.

Via

Access Paper or Ask Questions

Bias-Reduced Uncertainty Estimation for Deep Neural Classifiers

Sep 30, 2018

Yonatan Geifman, Guy Uziel, Ran El-Yaniv

Figure 1 for Bias-Reduced Uncertainty Estimation for Deep Neural Classifiers

Figure 2 for Bias-Reduced Uncertainty Estimation for Deep Neural Classifiers

Figure 3 for Bias-Reduced Uncertainty Estimation for Deep Neural Classifiers

Figure 4 for Bias-Reduced Uncertainty Estimation for Deep Neural Classifiers

Abstract:We consider the problem of uncertainty estimation in the context of (non-Bayesian) deep neural classification. In this context, all known methods are based on extracting uncertainty signals from a trained network optimized to solve the classification problem at hand. We demonstrate that such techniques tend to introduce biased estimates for instances whose predictions are supposed to be highly confident. We argue that this deficiency is an artifact of the dynamics of training with SGD-like optimizers, and it has some properties similar to overfitting. Based on this observation, we develop an uncertainty estimation algorithm that selectively estimates the uncertainty of highly confident points, using earlier snapshots of the trained model, before their estimates are jittered (and way before they are ready for actual classification). We present extensive experiments indicating that the proposed algorithm provides uncertainty estimates that are consistently better than all known methods.

Via

Access Paper or Ask Questions

Deep Active Learning over the Long Tail

Nov 02, 2017

Yonatan Geifman, Ran El-Yaniv

Figure 1 for Deep Active Learning over the Long Tail

Figure 2 for Deep Active Learning over the Long Tail

Figure 3 for Deep Active Learning over the Long Tail

Figure 4 for Deep Active Learning over the Long Tail

Abstract:This paper is concerned with pool-based active learning for deep neural networks. Motivated by coreset dataset compression ideas, we present a novel active learning algorithm that queries consecutive points from the pool using farthest-first traversals in the space of neural activation over a representation layer. We show consistent and overwhelming improvement in sample complexity over passive learning (random sampling) for three datasets: MNIST, CIFAR-10, and CIFAR-100. In addition, our algorithm outperforms the traditional uncertainty sampling technique (obtained using softmax activations), and we identify cases where uncertainty sampling is only slightly better than random sampling.

Via

Access Paper or Ask Questions

Selective Classification for Deep Neural Networks

Jun 01, 2017

Yonatan Geifman, Ran El-Yaniv

Figure 1 for Selective Classification for Deep Neural Networks

Figure 2 for Selective Classification for Deep Neural Networks

Figure 3 for Selective Classification for Deep Neural Networks

Figure 4 for Selective Classification for Deep Neural Networks

Abstract:Selective classification techniques (also known as reject option) have not yet been considered in the context of deep neural networks (DNNs). These techniques can potentially significantly improve DNNs prediction performance by trading-off coverage. In this paper we propose a method to construct a selective classifier given a trained neural network. Our method allows a user to set a desired risk level. At test time, the classifier rejects instances as needed, to grant the desired risk (with high probability). Empirical results over CIFAR and ImageNet convincingly demonstrate the viability of our method, which opens up possibilities to operate DNNs in mission-critical applications. For example, using our method an unprecedented 2% error in top-5 ImageNet classification can be guaranteed with probability 99.9%, and almost 60% test coverage.

Via

Access Paper or Ask Questions

The Prediction Advantage: A Universally Meaningful Performance Measure for Classification and Regression

May 26, 2017

Ran El-Yaniv, Yonatan Geifman, Yair Wiener

Figure 1 for The Prediction Advantage: A Universally Meaningful Performance Measure for Classification and Regression

Figure 2 for The Prediction Advantage: A Universally Meaningful Performance Measure for Classification and Regression

Abstract:We introduce the Prediction Advantage (PA), a novel performance measure for prediction functions under any loss function (e.g., classification or regression). The PA is defined as the performance advantage relative to the Bayesian risk restricted to knowing only the distribution of the labels. We derive the PA for well-known loss functions, including 0/1 loss, cross-entropy loss, absolute loss, and squared loss. In the latter case, the PA is identical to the well-known R-squared measure, widely used in statistics. The use of the PA ensures meaningful quantification of prediction performance, which is not guaranteed, for example, when dealing with noisy imbalanced classification problems. We argue that among several known alternative performance measures, PA is the best (and only) quantity ensuring meaningfulness for all noise and imbalance levels.

Via

Access Paper or Ask Questions