Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Dmitry Sorokin

BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Jun 14, 2024

Yuri Kuratov, Aydar Bulatov, Petr Anokhin, Ivan Rodkin, Dmitry Sorokin, Artyom Sorokin, Mikhail Burtsev

Figure 1 for BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Figure 2 for BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Figure 3 for BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Figure 4 for BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Abstract:In recent years, the input context sizes of large language models (LLMs) have increased dramatically. However, existing evaluation methods have not kept pace, failing to comprehensively assess the efficiency of models in handling long contexts. To bridge this gap, we introduce the BABILong benchmark, designed to test language models' ability to reason across facts distributed in extremely long documents. BABILong includes a diverse set of 20 reasoning tasks, including fact chaining, simple induction, deduction, counting, and handling lists/sets. These tasks are challenging on their own, and even more demanding when the required facts are scattered across long natural text. Our evaluations show that popular LLMs effectively utilize only 10-20\% of the context and their performance declines sharply with increased reasoning complexity. Among alternatives to in-context reasoning, Retrieval-Augmented Generation methods achieve a modest 60\% accuracy on single-fact question answering, independent of context length. Among context extension methods, the highest performance is demonstrated by recurrent memory transformers, enabling the processing of lengths up to 11 million tokens. The BABILong benchmark is extendable to any length to support the evaluation of new upcoming models with increased capabilities, and we provide splits up to 1 million token lengths.

Via

Access Paper or Ask Questions

In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss

Feb 21, 2024

Yuri Kuratov, Aydar Bulatov, Petr Anokhin, Dmitry Sorokin, Artyom Sorokin, Mikhail Burtsev

Figure 1 for In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss

Figure 2 for In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss

Figure 3 for In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss

Figure 4 for In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss

Abstract:This paper addresses the challenge of processing long documents using generative transformer models. To evaluate different approaches, we introduce BABILong, a new benchmark designed to assess model capabilities in extracting and processing distributed facts within extensive texts. Our evaluation, which includes benchmarks for GPT-4 and RAG, reveals that common methods are effective only for sequences up to $10^4$ elements. In contrast, fine-tuning GPT-2 with recurrent memory augmentations enables it to handle tasks involving up to $11\times 10^6$ elements. This achievement marks a substantial leap, as it is by far the longest input processed by any neural network model to date, demonstrating a significant improvement in the processing capabilities for long sequences.

* 11M tokens, fix qa3 min facts per task in Table 1

Via

Access Paper or Ask Questions

TreeDQN: Learning to minimize Branch-and-Bound tree

Jun 09, 2023

Dmitry Sorokin, Alexander Kostin

Abstract:Combinatorial optimization problems require an exhaustive search to find the optimal solution. A convenient approach to solving combinatorial optimization tasks in the form of Mixed Integer Linear Programs is Branch-and-Bound. Branch-and-Bound solver splits a task into two parts dividing the domain of an integer variable, then it solves them recursively, producing a tree of nested sub-tasks. The efficiency of the solver depends on the branchning heuristic used to select a variable for splitting. In the present work, we propose a reinforcement learning method that can efficiently learn the branching heuristic. We view the variable selection task as a tree Markov Decision Process, prove that the Bellman operator adapted for the tree Markov Decision Process is contracting in mean, and propose a modified learning objective for the reinforcement learning agent. Our agent requires less training data and produces smaller trees compared to previous reinforcement learning methods.

* Submitted to NeurIPS 2023

Via

Access Paper or Ask Questions

Insights From the NeurIPS 2021 NetHack Challenge

Mar 22, 2022

Eric Hambro, Sharada Mohanty, Dmitrii Babaev, Minwoo Byeon, Dipam Chakraborty, Edward Grefenstette, Minqi Jiang, Daejin Jo, Anssi Kanervisto, Jongmin Kim(+19 more)

Figure 1 for Insights From the NeurIPS 2021 NetHack Challenge

Figure 2 for Insights From the NeurIPS 2021 NetHack Challenge

Figure 3 for Insights From the NeurIPS 2021 NetHack Challenge

Figure 4 for Insights From the NeurIPS 2021 NetHack Challenge

Abstract:In this report, we summarize the takeaways from the first NeurIPS 2021 NetHack Challenge. Participants were tasked with developing a program or agent that can win (i.e., 'ascend' in) the popular dungeon-crawler game of NetHack by interacting with the NetHack Learning Environment (NLE), a scalable, procedurally generated, and challenging Gym environment for reinforcement learning (RL). The challenge showcased community-driven progress in AI with many diverse approaches significantly beating the previously best results on NetHack. Furthermore, it served as a direct comparison between neural (e.g., deep RL) and symbolic AI, as well as hybrid systems, demonstrating that on NetHack symbolic bots currently outperform deep RL by a large margin. Lastly, no agent got close to winning the game, illustrating NetHack's suitability as a long-term benchmark for AI research.

* Under review at PMLR for the NeuRIPS 2021 Competition Workshop Track, 10 pages + 10 in appendices

Via

Access Paper or Ask Questions

Aligning an optical interferometer with beam divergence control and continuous action space

Jul 09, 2021

Stepan Makarenko, Dmitry Sorokin, Alexander Ulanov, A. I. Lvovsky

Figure 1 for Aligning an optical interferometer with beam divergence control and continuous action space

Figure 2 for Aligning an optical interferometer with beam divergence control and continuous action space

Figure 3 for Aligning an optical interferometer with beam divergence control and continuous action space

Figure 4 for Aligning an optical interferometer with beam divergence control and continuous action space

Abstract:Reinforcement learning is finding its way to real-world problem application, transferring from simulated environments to physical setups. In this work, we implement vision-based alignment of an optical Mach-Zehnder interferometer with a confocal telescope in one arm, which controls the diameter and divergence of the corresponding beam. We use a continuous action space; exponential scaling enables us to handle actions within a range of over two orders of magnitude. Our agent trains only in a simulated environment with domain randomizations. In an experimental evaluation, the agent significantly outperforms an existing solution and a human expert.

* 12 pages, 5 figures

Via

Access Paper or Ask Questions

Adaptation of Quadruped Robot Locomotion with Meta-Learning

Jul 08, 2021

Arsen Kuzhamuratov, Dmitry Sorokin, Alexander Ulanov, A. I. Lvovsky

Figure 1 for Adaptation of Quadruped Robot Locomotion with Meta-Learning

Figure 2 for Adaptation of Quadruped Robot Locomotion with Meta-Learning

Figure 3 for Adaptation of Quadruped Robot Locomotion with Meta-Learning

Figure 4 for Adaptation of Quadruped Robot Locomotion with Meta-Learning

Abstract:Animals have remarkable abilities to adapt locomotion to different terrains and tasks. However, robots trained by means of reinforcement learning are typically able to solve only a single task and a transferred policy is usually inferior to that trained from scratch. In this work, we demonstrate that meta-reinforcement learning can be used to successfully train a robot capable to solve a wide range of locomotion tasks. The performance of the meta-trained robot is similar to that of a robot that is trained on a single task.

* 14 pages, 6 figures

Via

Access Paper or Ask Questions

Interferobot: aligning an optical interferometer by a reinforcement learning agent

Jun 03, 2020

Dmitry Sorokin, Alexander Ulanov, Ekaterina Sazhina, Alexander Lvovsky

Figure 1 for Interferobot: aligning an optical interferometer by a reinforcement learning agent

Figure 2 for Interferobot: aligning an optical interferometer by a reinforcement learning agent

Figure 3 for Interferobot: aligning an optical interferometer by a reinforcement learning agent

Figure 4 for Interferobot: aligning an optical interferometer by a reinforcement learning agent

Abstract:Limitations in acquiring training data restrict potential applications of deep reinforcement learning (RL) methods to the training of real-world robots. Here we train an RL agent to align a Mach-Zehnder interferometer, which is an essential part of many optical experiments, based on images of interference fringes acquired by a monocular camera. The agent is trained in a simulated environment, without any hand-coded features or a priori information about the physics, and subsequently transferred to a physical interferometer. Thanks to a set of domain randomizations simulating uncertainties in physical measurements, the agent successfully aligns this interferometer without any fine tuning, achieving a performance level of a human expert.

Via

Access Paper or Ask Questions