Team Ai
Datasetpublic

BEE-spoke-data/code-tutorials-en

Dataset Card for "code-tutorials-en" en only 100 words or more reading ease of 50 or more DatasetDict({ train: Dataset({ features: ['text', 'url', 'dump', 'source', 'word_count', 'flesch_reading_ease'], num_rows: 223162 }) validation: Dataset({ features: ['text', 'url', 'dump', 'source', 'word_count', 'flesch_reading_ease'], num_rows: 5873 }) test: Dataset({ features: ['text', 'url', 'dump', 'source', 'word_count'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code-tutorials-en.

sourceHugging Faceodc-byupdated 10mo agoView on Hugging Face
1likes134downloads
README.md92 linesDownload Raw Back to root
1---2configs:3- config_name: default4  data_files:5  - split: train6    path: data/train-*7  - split: validation8    path: data/validation-*9  - split: test10    path: data/test-*11- config_name: unfiltered12  data_files:13  - split: train14    path: unfiltered/train-*15dataset_info:16- config_name: default17  features:18  - name: text19    dtype: string20  - name: url21    dtype: string22  - name: dump23    dtype: string24  - name: source25    dtype: string26  - name: word_count27    dtype: int6428  - name: flesch_reading_ease29    dtype: float6430  splits:31  - name: train32    num_bytes: 2003343392.865814233    num_examples: 22316234  - name: validation35    num_bytes: 52722397.837897736    num_examples: 587337  - name: test38    num_bytes: 52722397.837897739    num_examples: 587340  download_size: 113745702741  dataset_size: 2108788188.541609842- config_name: unfiltered43  features:44  - name: text45    dtype: string46  - name: url47    dtype: string48  - name: dump49    dtype: string50  - name: source51    dtype: string52  - name: word_count53    dtype: int6454  - name: flesch_reading_ease55    dtype: float6456  splits:57  - name: train58    num_bytes: 345299837259    num_examples: 38464660  download_size: 185937582461  dataset_size: 345299837262source_datasets: mponty/code_tutorials63license: odc-by64task_categories:65- text-generation66language:67- en68size_categories:69- 100K<n<1M70---71# Dataset Card for "code-tutorials-en"72 73- `en` only74- 100 words or more75- reading ease of 50 or more76 77```78DatasetDict({79    train: Dataset({80        features: ['text', 'url', 'dump', 'source', 'word_count', 'flesch_reading_ease'],81        num_rows: 22316282    })83    validation: Dataset({84        features: ['text', 'url', 'dump', 'source', 'word_count', 'flesch_reading_ease'],85        num_rows: 587386    })87    test: Dataset({88        features: ['text', 'url', 'dump', 'source', 'word_count', 'flesch_reading_ease'],89        num_rows: 587390    })91})92```