Paloma: A Benchmark for Evaluating Language Model Fit
Abstract
Language models (LMs) commonly report perplexity on monolithic data held out from training. Implicitly or explicitly, this data is composed of domainsx2013varying distributions of language. Rather than assuming perplexity on one distribution extrapolates to others, Perplexity Analysis for Language Model Assessment (Paloma), measures LM fit to 585 text domains, ranging from nytimes.com to r/depression on Reddit. We invite submissions to our benchmark and organize results by comparability based on compliance with guidelines such as removal of benchmark contamination from pretraining. Submissions can also record parameter and training token count to make comparisons of Pareto efficiency for performance as a function of these measures of cost. We populate our benchmark with results from 6 baselines pretrained on popular corpora. In case studies, we demonstrate analyses that are possible with Paloma, such as finding that pretraining without data beyond Common Crawl leads to inconsistent fit to many domains.
Community
The dataset is available on the hub allenai/paloma. If you add the arxiv.org/abs/2312.10523
somewhere in the README's citation section the paper and dataset should get automatically linked :)
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- MindLLM: Pre-training Lightweight Large Language Model from Scratch, Evaluations and Domain Applications (2023)
- When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource Languages (2023)
- FinGPT: Large Generative Models for a Small Language (2023)
- Efficiently Adapting Pretrained Language Models To New Languages (2023)
- Data Similarity is Not Enough to Explain Language Model Performance (2023)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper