Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:MESS+: Energy-Optimal Inferencing in Language Model Zoos with Service Level Guarantees

Oct 31, 2024

Ryan Zhang, Herbert Woisetschläger, Shiqiang Wang, Hans Arno Jacobsen

Figure 1 for MESS+: Energy-Optimal Inferencing in Language Model Zoos with Service Level Guarantees

Figure 2 for MESS+: Energy-Optimal Inferencing in Language Model Zoos with Service Level Guarantees

Figure 3 for MESS+: Energy-Optimal Inferencing in Language Model Zoos with Service Level Guarantees

Figure 4 for MESS+: Energy-Optimal Inferencing in Language Model Zoos with Service Level Guarantees

Share this with someone who'll enjoy it:

Abstract:Open-weight large language model (LLM) zoos allow users to quickly integrate state-of-the-art models into systems. Despite increasing availability, selecting the most appropriate model for a given task still largely relies on public benchmark leaderboards and educated guesses. This can be unsatisfactory for both inference service providers and end users, where the providers usually prioritize cost efficiency, while the end users usually prioritize model output quality for their inference requests. In commercial settings, these two priorities are often brought together in Service Level Agreements (SLA). We present MESS+, an online stochastic optimization algorithm for energy-optimal model selection from a model zoo, which works on a per-inference-request basis. For a given SLA that requires high accuracy, we are up to 2.5x more energy efficient with MESS+ than with randomly selecting an LLM from the zoo while maintaining SLA quality constraints.

* Accepted at the 2024 Workshop on Adaptive Foundation Models in conjunction with NeurIPS 2024

View paper on

Share this with someone who'll enjoy it:

Title:MESS+: Energy-Optimal Inferencing in Language Model Zoos with Service Level Guarantees

Paper and Code