Team Ai
Datasetpublic

HuggingFaceM4/OBELICS

Dataset Card for OBELICS OBELICS is an open, massive, and curated collection of interleaved image-text web documents, containing 141M English documents, 115B text tokens, and 353M images, extracted from Common Crawl dumps between February 2020 and February 2023. The collection and filtering steps are described in our paper. Interleaved image-text web documents are a succession of text paragraphs interleaved by images, such as web pages that contain images. Models trained on… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceM4/OBELICS.

sourceHugging Facecc-by-4.0updated 3y agoView on Hugging Face
174likes10kdownloads
README.md155 linesDownload Raw Back to root
1---2language:3- en4license: cc-by-4.05size_categories:6- 100M<n<1B7pretty_name: OBELICS8configs:9- config_name: default10  data_files:11  - split: train12    path: data/train-*13- config_name: opt_out_docs_removed_2023_07_1214  data_files:15  - split: train16    path: opt_out_docs_removed_2023_07_12/train-*17dataset_info:18- config_name: default19  features:20  - name: images21    sequence: string22  - name: metadata23    dtype: string24  - name: general_metadata25    dtype: string26  - name: texts27    sequence: string28  splits:29  - name: train30    num_bytes: 71572471719231    num_examples: 14104769732  download_size: 7152062965533  dataset_size: 71572471719234- config_name: opt_out_docs_removed_2023_07_1235  features:36  - name: images37    sequence: string38  - name: metadata39    dtype: string40  - name: general_metadata41    dtype: string42  - name: texts43    sequence: string44  splits:45  - name: train46    num_bytes: 68463831421547    num_examples: 13464885548  download_size: 26650109292049  dataset_size: 68463831421550---51# Dataset Card for OBELICS52 53## Dataset Description54 55- **Visualization of OBELICS web documents:** https://huggingface.co/spaces/HuggingFaceM4/obelics_visualization56- **Paper:** [OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents](https://arxiv.org/abs/2306.16527)57- **Repository:** https://github.com/huggingface/OBELICS58- **Point of Contact: hugo@huggingface.co**59 60`OBELICS` is an open, massive, and curated collection of interleaved image-text web documents, containing 141M English documents, 115B text tokens, and 353M images, extracted from Common Crawl dumps between February 2020 and February 2023. The collection and filtering steps are described in our [paper](https://huggingface.co/papers/2306.16527).61 62Interleaved image-text web documents are a succession of text paragraphs interleaved by images, such as web pages that contain images. Models trained on these web documents outperform vision and language models trained solely on image-text pairs on various benchmarks. They can also generate long and coherent text about a set of multiple images. As an example, we trained [IDEFICS](https://huggingface.co/HuggingFaceM4/idefics-80b), a visual language model that accepts arbitrary sequences of image and text inputs and produces text outputs.63 64We provide an [interactive visualization](https://atlas.nomic.ai/map/f2fba2aa-3647-4f49-a0f3-9347daeee499/ee4a84bd-f125-4bcc-a683-1b4e231cb10f) of OBELICS that allows exploring the content of OBELICS. The map shows a subset of 11M of the 141M documents.65 66[![OBELICS Nomic map](assets/nomic_map.png)](https://atlas.nomic.ai/map/f2fba2aa-3647-4f49-a0f3-9347daeee499/ee4a84bd-f125-4bcc-a683-1b4e231cb10f)67 68 69## Data Fields70 71An example of a sample looks as follows:72```73# The example has been cropped74 75{76    'images': [77        'https://cdn.motor1.com/images/mgl/oRKO0/s1/lamborghini-urus-original-carbon-fiber-accessories.jpg',78        None79    ],80    'metadata': '[{"document_url": "https://lamborghinichat.com/forum/news/vw-group-allegedly-receives-offer-to-sell-lamborghini-for-9-2-billion.728/", "unformatted_src": "https://cdn.motor1.com/images/mgl/oRKO0/s1/lamborghini-urus-original-carbon-fiber-accessories.jpg", "src": "https://cdn.motor1.com/images/mgl/oRKO0/s1/lamborghini-urus-original-carbon-fiber-accessories.jpg", "formatted_filename": "lamborghini urus original carbon fiber accessories", "alt_text": "VW Group Allegedly Receives Offer To Sell Lamborghini For $9.2 Billion", "original_width": 1920, "original_height": 1080, "format": "jpeg"}, null]',81    'general_metadata': '{"url": "https://lamborghinichat.com/forum/news/vw-group-allegedly-receives-offer-to-sell-lamborghini-for-9-2-billion.728/", "warc_filename": "crawl-data/CC-MAIN-2021-25/segments/1623488528979.69/warc/CC-MAIN-20210623011557-20210623041557-00312.warc.gz", "warc_record_offset": 322560850, "warc_record_length": 17143}',82    'texts': [83        None,84        'The buyer would get everything, including Lambo\'s headquarters.\n\nThe investment groupQuantum Group AG has submitted a€7.5 billion ($9.2 billion at current exchange rates) offer to purchase Lamborghini from Volkswagen Group, Autocar reports. There\'s no info yet about whether VW intends to accept the offer or further negotiate the deal.\n\nQuantum ... Group Chief Executive Herbert Diess said at the time.'85    ]86}87```88 89Each sample is composed of the same 4 fields: `images`, `texts`, `metadata`, and `general_metadata`. `images` and `texts` are two lists of the same size, where for each index, one element and only one is not `None`. For example, for the interleaved web document `<image_1>text<image_2>`, we would find `[image_1, None, image_2]` in `images` and `[None, text, None]` in `texts`.90 91The images are replaced by their URLs, and the users need to download the images, for instance, with the library [img2dataset](https://github.com/rom1504/img2dataset).92 93`metadata` is the string representation of a list containing information about each of the images. It has the same length as `texts` and `images` and logs for each image relevant information such as original source document, unformatted source, alternative text if present, etc.94 95`general_metadata` is the string representation of a dictionary containing the URL of the document, and information regarding the extraction from Common Crawl snapshots.96 97## Size and Data Splits98 99There is only one split, `train`, that contains 141,047,697 documents.100 101`OBELICS` with images replaced by their URLs weighs 666.6 GB (😈) in arrow format and 377 GB in the uploaded `parquet` format.102 103 104## Considerations for Using the Data105 106 107### Discussion of Biases108 109A subset of this dataset `train`, of ~50k was evaluated using the Data Measurements Tool, with a particular focus on the nPMI metric 110> nPMI scores for a word help to identify potentially problematic associations, ranked by how close the association is.111> nPMI bias scores for paired words help to identify how word associations are skewed between the selected selected words (Aka et al., 2021).112> You can select from gender and sexual orientation identity terms that appear in the dataset at least 10 times.113> The resulting ranked words are those that co-occur with both identity terms.114> The more positive the score, the more associated the word is with the first identity term. The more negative the score, the more associated the word is with the second identity term.115 116While there was a positive skew of words relating occupations e.g _`government`_, _`jobs`_ towards she, her, and similar attributions of the masculine and feminine words to they and them, more harmful words attributions such as _`escort`_ and even _`colour`_ presented with greater attributions to she, her and him, his, respectively.117 118![Data Measurement Tool Associations Eval](assets/DMT_eval.png)119 120We welcome users to explore the [Data Measurements nPMI Visualitons for OBELICS](https://huggingface.co/spaces/HuggingFaceM4/IDEFICS_Data_Measurement_Tool) further and to see the [idefics-9b model card](https://huggingface.co/HuggingFaceM4/idefics-9b) for further Bias considerations.121 122## Opted-out content123 124To respect the preferences of content creators, we removed from OBELICS all images for which creators explicitly opted out of AI model training. We used the [Spawning API](https://api.spawning.ai/spawning-api) to verify that the images in the dataset respect the original copyright owners’ choices.125 126However, due to an error on our side, we did not remove entire documents (i.e., URLs) that opted out of AI model training. As of July 12, 2023, it represents 4.25% of the totality of OBELICS. The config `opt_out_docs_removed_2023_07_12` applies the correct filtering at the web document level as of July 2023: `ds = load_dataset("HuggingFaceM4/OBELICS", "opt_out_docs_removed_2023_07_12")`.127 128We recommend users of OBELICS to regularly check every document against the API.129 130## Content warnings131 132Despite our efforts in filtering, OBELICS contains a small proportion of documents that are not suitable for all audiences. For instance, while navigating the interactive map, you might find the cluster named "Sex" which predominantly contains descriptions of pornographic movies along with pornographic images. Other clusters would contain advertising for sex workers or reports of violent shootings. In our experience, these documents represent a small proportion of all the documents.133 134## Terms of Use135 136By using the dataset, you agree to comply with the original licenses of the source content as well as the dataset license (CC-BY-4.0). Additionally, if you use this dataset to train a Machine Learning model, you agree to disclose your use of the dataset when releasing the model or an ML application using the model.137 138### Licensing Information139 140License CC-BY-4.0.141 142### Citation Information143 144If you are using this dataset, please cite145```146@misc{laurencon2023obelics,147      title={OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents},148      author={Hugo Laurençon and Lucile Saulnier and Léo Tronchon and Stas Bekman and Amanpreet Singh and Anton Lozhkov and Thomas Wang and Siddharth Karamcheti and Alexander M. Rush and Douwe Kiela and Matthieu Cord and Victor Sanh},149      year={2023},150      eprint={2306.16527},151      archivePrefix={arXiv},152      primaryClass={cs.IR}153}154```155