Chaiverse calculates the ELO of LLMs via human feedback, obtained through blind selection testing using the Chai app. This involves gathering positive and negative upvotes from randomly selected models on our leaderboard, as experienced by end users. With over one million samples a day, we are able to build a ranking of LLMs - which for our use-case, seems to be the most accurate.

