Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:No Free Lunch From Random Feature Ensembles

Dec 06, 2024

Benjamin S. Ruben, William L. Tong, Hamza Tahir Chaudhry, Cengiz Pehlevan

Figure 1 for No Free Lunch From Random Feature Ensembles

Figure 2 for No Free Lunch From Random Feature Ensembles

Figure 3 for No Free Lunch From Random Feature Ensembles

Figure 4 for No Free Lunch From Random Feature Ensembles

Share this with someone who'll enjoy it:

Abstract:Given a budget on total model size, one must decide whether to train a single, large neural network or to combine the predictions of many smaller networks. We study this trade-off for ensembles of random-feature ridge regression models. We prove that when a fixed number of trainable parameters are partitioned among $K$ independently trained models, $K=1$ achieves optimal performance, provided the ridge parameter is optimally tuned. We then derive scaling laws which describe how the test risk of an ensemble of regression models decays with its total size. We identify conditions on the kernel and task eigenstructure under which ensembles can achieve near-optimal scaling laws. Training ensembles of deep convolutional neural networks on CIFAR-10 and a transformer architecture on C4, we find that a single large network outperforms any ensemble of networks with the same total number of parameters, provided the weight decay and feature-learning strength are tuned to their optimal values.

View paper on

Share this with someone who'll enjoy it:

Title:No Free Lunch From Random Feature Ensembles

Paper and Code