Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Kening Zhang

Neural Program Synthesis By Self-Learning

Oct 13, 2019

Yifan Xu, Lu Dai, Udaikaran Singh, Kening Zhang, Zhuowen Tu

Figure 1 for Neural Program Synthesis By Self-Learning

Figure 2 for Neural Program Synthesis By Self-Learning

Figure 3 for Neural Program Synthesis By Self-Learning

Abstract:Neural inductive program synthesis is a task generating instructions that can produce desired outputs from given inputs. In this paper, we focus on the generation of a chunk of assembly code that can be executed to match a state change inside the CPU and RAM. We develop a neural program synthesis algorithm, AutoAssemblet, learned via self-learning reinforcement learning that explores the large code space efficiently. Policy networks and value networks are learned to reduce the breadth and depth of the Monte Carlo Tree Search, resulting in better synthesis performance. We also propose an effective multi-entropy policy sampling technique to alleviate online update correlations. We apply AutoAssemblet to basic programming tasks and show significant higher success rates compared to several competing baselines.

Via

Access Paper or Ask Questions

Rethinking Exposure Bias In Language Modeling

Oct 13, 2019

Yifan Xu, Kening Zhang, Haoyu Dong, Yuezhou Sun, Wenlong Zhao, Zhuowen Tu

Figure 1 for Rethinking Exposure Bias In Language Modeling

Figure 2 for Rethinking Exposure Bias In Language Modeling

Figure 3 for Rethinking Exposure Bias In Language Modeling

Figure 4 for Rethinking Exposure Bias In Language Modeling

Abstract:Exposure bias describes the phenomenon that a language model trained under the teacher forcing schema may perform poorly at the inference stage when its predictions are conditioned on its previous predictions unseen from the training corpus. Recently, several generative adversarial networks (GANs) and reinforcement learning (RL) methods have been introduced to alleviate this problem. Nonetheless, a common issue in RL and GANs training is the sparsity of reward signals. In this paper, we adopt two simple strategies, multi-range reinforcing, and multi-entropy sampling, to amplify and denoise the reward signal. Our model produces an improvement over competing models with regards to BLEU scores and road exam, a new metric we designed to measure the robustness against exposure bias in language models.

Via

Access Paper or Ask Questions