davidschulte/ESM_AmazonScience__massive_ru-RU
013
1---2base_model: bert-base-multilingual-uncased3datasets:4- AmazonScience/massive5license: apache-2.06tags:7- embedding_space_map8- BaseLM:bert-base-multilingual-uncased9---10 11# ESM AmazonScience/massive12 13<!-- Provide a quick summary of what the model is/does. -->14 15 16 17## Model Details18 19### Model Description20 21<!-- Provide a longer summary of what this model is. -->22 23ESM24 25- **Developed by:** David Schulte26- **Model type:** ESM27- **Base Model:** bert-base-multilingual-uncased28- **Intermediate Task:** AmazonScience/massive29- **ESM architecture:** linear30- **ESM embedding dimension:** 76831- **Language(s) (NLP):** [More Information Needed]32- **License:** Apache-2.0 license33- **ESM version:** 0.1.034 35## Training Details36 37### Intermediate Task38- **Task ID:** AmazonScience/massive39- **Subset [optional]:** ru-RU40- **Text Column:** annot_utt41- **Label Column:** scenario42- **Dataset Split:** train43- **Sample size [optional]:** 1000044- **Sample seed [optional]:** 4245 46### Training Procedure [optional]47 48<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->49 50#### Language Model Training Hyperparameters [optional]51- **Epochs:** 352- **Batch size:** 3253- **Learning rate:** 2e-0554- **Weight Decay:** 0.0155- **Optimizer**: AdamW56 57### ESM Training Hyperparameters [optional]58- **Epochs:** 1059- **Batch size:** 3260- **Learning rate:** 0.00161- **Weight Decay:** 0.0162- **Optimizer**: AdamW63 64 65### Additional trainiung details [optional]66 67 68## Model evaluation69 70### Evaluation of fine-tuned language model [optional]71 72 73### Evaluation of ESM [optional]74MSE: 75 76### Additional evaluation details [optional]77 78 79## What are Embedding Space Maps used for?80Embedding Space Maps are a part of ESM-LogME, a efficient method for finding intermediate datasets for transfer learning. There are two reasons to use ESM-LogME:81 82### You don't have enough training data for your problem83If you don't have a enough training data for your problem, just use ESM-LogME to find more.84You can supplement model training by including publicly available datasets in the training process. 85 861. Fine-tune a language model on suitable intermediate dataset.872. Fine-tune the resulting model on your target dataset.88 89This workflow is called intermediate task transfer learning and it can significantly improve the target performance.90 91But what is a suitable dataset for your problem? ESM-LogME enable you to quickly rank thousands of datasets on the Hugging Face Hub by how well they are exptected to transfer to your target task.92 93### You want to find similar datasets to your target dataset94Using ESM-LogME can be used like search engine on the Hugging Face Hub. You can find similar tasks to your target task without having to rely on heuristics. ESM-LogME estimates how language models fine-tuned on each intermediate task would benefinit your target task. This quantitative approach combines the effects of domain similarity and task similarity. 95 96## How can I use ESM-LogME / ESMs?97[](https://pypi.org/project/hf-dataset-selector)98 99We release **hf-dataset-selector**, a Python package for intermediate task selection using Embedding Space Maps.100 101**hf-dataset-selector** fetches ESMs for a given language model and uses it to find the best dataset for applying intermediate training to the target task. ESMs are found by their tags on the Huggingface Hub.102 103```python104from hfselect import Dataset, compute_task_ranking105 106# Load target dataset from the Hugging Face Hub107dataset = Dataset.from_hugging_face(108 name="stanfordnlp/imdb",109 split="train",110 text_col="text",111 label_col="label",112 is_regression=False,113 num_examples=1000,114 seed=42115)116 117# Fetch ESMs and rank tasks118task_ranking = compute_task_ranking(119 dataset=dataset,120 model_name="bert-base-multilingual-uncased"121)122 123# Display top 5 recommendations124print(task_ranking[:5])125```126```python1271. davanstrien/test_imdb_embedd2 Score: -0.6185291282. davanstrien/test_imdb_embedd Score: -0.6186441293. davanstrien/test1 Score: -0.6193341304. stanfordnlp/imdb Score: -0.6194541315. stanfordnlp/sst Score: -0.62995132```133 134| Rank | Task ID | Task Subset | Text Column | Label Column | Task Split | Num Examples | ESM Architecture | Score |135|-------:|:------------------------------|:----------------|:--------------|:---------------|:-------------|---------------:|:-------------------|----------:|136| 1 | davanstrien/test_imdb_embedd2 | default | text | label | train | 10000 | linear | -0.618529 |137| 2 | davanstrien/test_imdb_embedd | default | text | label | train | 10000 | linear | -0.618644 |138| 3 | davanstrien/test1 | default | text | label | train | 10000 | linear | -0.619334 |139| 4 | stanfordnlp/imdb | plain_text | text | label | train | 10000 | linear | -0.619454 |140| 5 | stanfordnlp/sst | dictionary | phrase | label | dictionary | 10000 | linear | -0.62995 |141| 6 | stanfordnlp/sst | default | sentence | label | train | 8544 | linear | -0.63312 |142| 7 | kuroneko5943/snap21 | CDs_and_Vinyl_5 | sentence | label | train | 6974 | linear | -0.634365 |143| 8 | kuroneko5943/snap21 | Video_Games_5 | sentence | label | train | 6997 | linear | -0.638787 |144| 9 | kuroneko5943/snap21 | Movies_and_TV_5 | sentence | label | train | 6989 | linear | -0.639068 |145| 10 | fancyzhx/amazon_polarity | amazon_polarity | content | label | train | 10000 | linear | -0.639718 |146 147For more information on how to use ESMs please have a look at the [official Github repository](https://github.com/davidschulte/hf-dataset-selector). We provide documentation further documentation and tutorials for finding intermediate datasets and training your own ESMs.148 149 150## How do Embedding Space Maps work?151 152<!-- This section describes the evaluation protocols and provides the results. -->153Embedding Space Maps (ESMs) are neural networks that approximate the effect of fine-tuning a language model on a task. They can be used to quickly transform embeddings from a base model to approximate how a fine-tuned model would embed the the input text.154ESMs can be used for intermediate task selection with the ESM-LogME workflow.155 156## How can I use Embedding Space Maps for Intermediate Task Selection?157 158## Citation159 160 161<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->162If you are using this Embedding Space Maps, please cite our [paper](https://aclanthology.org/2024.emnlp-main.529/).163 164**BibTeX:**165 166 167```168@inproceedings{schulte-etal-2024-less,169 title = "Less is More: Parameter-Efficient Selection of Intermediate Tasks for Transfer Learning",170 author = "Schulte, David and171 Hamborg, Felix and172 Akbik, Alan",173 editor = "Al-Onaizan, Yaser and174 Bansal, Mohit and175 Chen, Yun-Nung",176 booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",177 month = nov,178 year = "2024",179 address = "Miami, Florida, USA",180 publisher = "Association for Computational Linguistics",181 url = "https://aclanthology.org/2024.emnlp-main.529/",182 doi = "10.18653/v1/2024.emnlp-main.529",183 pages = "9431--9442",184 abstract = "Intermediate task transfer learning can greatly improve model performance. If, for example, one has little training data for emotion detection, first fine-tuning a language model on a sentiment classification dataset may improve performance strongly. But which task to choose for transfer learning? Prior methods producing useful task rankings are infeasible for large source pools, as they require forward passes through all source language models. We overcome this by introducing Embedding Space Maps (ESMs), light-weight neural networks that approximate the effect of fine-tuning a language model. We conduct the largest study on NLP task transferability and task selection with 12k source-target pairs. We find that applying ESMs on a prior method reduces execution time and disk space usage by factors of 10 and 278, respectively, while retaining high selection performance (avg. regret@5 score of 2.95)."185}186```187 188 189**APA:**190 191```192Schulte, D., Hamborg, F., & Akbik, A. (2024, November). Less is More: Parameter-Efficient Selection of Intermediate Tasks for Transfer Learning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 9431-9442).193```194 195## Additional Information196 197 