BEE-spoke-data/code-tutorials-en
Dataset Card for "code-tutorials-en" en only 100 words or more reading ease of 50 or more DatasetDict({ train: Dataset({ features: ['text', 'url', 'dump', 'source', 'word_count', 'flesch_reading_ease'], num_rows: 223162 }) validation: Dataset({ features: ['text', 'url', 'dump', 'source', 'word_count', 'flesch_reading_ease'], num_rows: 5873 }) test: Dataset({ features: ['text', 'url', 'dump', 'source', 'word_count'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code-tutorials-en.
1134
1---2configs:3- config_name: default4 data_files:5 - split: train6 path: data/train-*7 - split: validation8 path: data/validation-*9 - split: test10 path: data/test-*11- config_name: unfiltered12 data_files:13 - split: train14 path: unfiltered/train-*15dataset_info:16- config_name: default17 features:18 - name: text19 dtype: string20 - name: url21 dtype: string22 - name: dump23 dtype: string24 - name: source25 dtype: string26 - name: word_count27 dtype: int6428 - name: flesch_reading_ease29 dtype: float6430 splits:31 - name: train32 num_bytes: 2003343392.865814233 num_examples: 22316234 - name: validation35 num_bytes: 52722397.837897736 num_examples: 587337 - name: test38 num_bytes: 52722397.837897739 num_examples: 587340 download_size: 113745702741 dataset_size: 2108788188.541609842- config_name: unfiltered43 features:44 - name: text45 dtype: string46 - name: url47 dtype: string48 - name: dump49 dtype: string50 - name: source51 dtype: string52 - name: word_count53 dtype: int6454 - name: flesch_reading_ease55 dtype: float6456 splits:57 - name: train58 num_bytes: 345299837259 num_examples: 38464660 download_size: 185937582461 dataset_size: 345299837262source_datasets: mponty/code_tutorials63license: odc-by64task_categories:65- text-generation66language:67- en68size_categories:69- 100K<n<1M70---71# Dataset Card for "code-tutorials-en"72 73- `en` only74- 100 words or more75- reading ease of 50 or more76 77```78DatasetDict({79 train: Dataset({80 features: ['text', 'url', 'dump', 'source', 'word_count', 'flesch_reading_ease'],81 num_rows: 22316282 })83 validation: Dataset({84 features: ['text', 'url', 'dump', 'source', 'word_count', 'flesch_reading_ease'],85 num_rows: 587386 })87 test: Dataset({88 features: ['text', 'url', 'dump', 'source', 'word_count', 'flesch_reading_ease'],89 num_rows: 587390 })91})92```