
Lmsys has announced a significant update to its Chatbot Arena, introducing style control in its regression model. This development aims to separate the impact of style from substance in chatbot responses. The update, which involves adding style as a feature in logistic regression, is expected to address concerns that large language models (LLMs) are being manipulated to favor lengthy and well-formatted responses. Collaborators on this project include researchers Li Tianleli and Angelo Polous, who have utilized statistical techniques to decouple style from substance. The next phase of the project will focus on causal inference, with an invitation for causal experts to contribute. Observations suggest that major players like OpenAI and Google have been prominent in this 'style hacking' phenomenon, raising questions about the integrity of human evaluation benchmarks in AI.
