Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of Experts

May 22, 2024

Huy Nguyen, Nhat Ho, Alessandro Rinaldo

Figure 1 for Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of Experts

Figure 2 for Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of Experts

Figure 3 for Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of Experts

Figure 4 for Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of Experts

Share this with someone who'll enjoy it:

Abstract:The softmax gating function is arguably the most popular choice in mixture of experts modeling. Despite its widespread use in practice, softmax gating may lead to unnecessary competition among experts, potentially causing the undesirable phenomenon of representation collapse due to its inherent structure. In response, the sigmoid gating function has been recently proposed as an alternative and has been demonstrated empirically to achieve superior performance. However, a rigorous examination of the sigmoid gating function is lacking in current literature. In this paper, we verify theoretically that sigmoid gating, in fact, enjoys a higher sample efficiency than softmax gating for the statistical task of expert estimation. Towards that goal, we consider a regression framework in which the unknown regression function is modeled as a mixture of experts, and study the rates of convergence of the least squares estimator in the over-specified case in which the number of experts fitted is larger than the true value. We show that two gating regimes naturally arise and, in each of them, we formulate identifiability conditions for the expert functions and derive the corresponding convergence rates. In both cases, we find that experts formulated as feed-forward networks with commonly used activation such as $\mathrm{ReLU}$ and $\mathrm{GELU}$ enjoy faster convergence rates under sigmoid gating than softmax gating. Furthermore, given the same choice of experts, we demonstrate that the sigmoid gating function requires a smaller sample size than its softmax counterpart to attain the same error of expert estimation and, therefore, is more sample efficient.

* 31 pages, 2 figures. arXiv admin note: text overlap with arXiv:2402.02952

View paper on

Share this with someone who'll enjoy it:

Title:Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of Experts

Paper and Code