datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FineFineWeb-fasttext-seeddata
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-fasttext-seeddata.coconot_fasttext_filter_25B_tokensfasttext-datab2_code_fasttext_pos_ioi_neg_sql_eval_636d
mlfoundations-dev/b2_code_fasttext_pos_ioi_neg_sql_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
16.7
57.8
76.2
27.8
39.4
43.4
41.9
14.9
17.7
AIME24
Average Accuracy: 16.67% ± 1.33%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
10.00%
3
30
2
13.33%
4
30
3
16.67%
5
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_code_fasttext_pos_ioi_neg_sql_eval_636d.mc4-und-fasttexte1_code_fasttext_r1_eval_636d
mlfoundations-dev/e1_code_fasttext_r1_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
21.7
65.5
75.8
0.4
42.7
40.9
22.7
14.9
20.0
AIME24
Average Accuracy: 21.67% ± 2.07%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
13.33%
4
30
2
30.00%
9
30
3
16.67%
5
30
4
26.67%
8… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/e1_code_fasttext_r1_eval_636d.e1_code_fasttext_phi_temp2dclm-baseline-fasttextfineweb-edu-v2-fasttexttrain_fasttext_classifier_seed_code_worst_1distemist-fasttext-8-nerhttps://temu.bsc.es/multicardioner/chinese-japanese-classification-fasttext
Chinese-Japanese Text Classification Dataset
A clean, multi-label text dataset designed for training language and script classifiers (such as FastText) to accurately distinguish between closely related Chinese variants and Japanese.
Motivation & Background
Distinguishing between Chinese scripts and Japanese can be notoriously difficult for standard language detectors. This is primarily because Japanese incorporates Hanzi/Kanji (Chinese characters), which… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/chinese-japanese-classification-fasttext.e1_code_fasttext_r1_eval_2e29amazonreviews-fasttextdrugtemist-fasttext-85-nerhttps://temu.bsc.es/multicardioner/processed_arabic_embeddings_fasttext_ar_vectorsinstruction_filtering_scale_up_code_base_fasttext_per_domainb2_code_fasttext_pos_codeforces_neg_all_1k_eval_636d
mlfoundations-dev/b2_code_fasttext_pos_codeforces_neg_all_1k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
16.0
55.0
71.6
26.0
37.7
36.4
29.5
7.2
8.8
AIME24
Average Accuracy: 16.00% ± 1.32%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
16.67%
5
30
2
16.67%
5
30
3… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_code_fasttext_pos_codeforces_neg_all_1k_eval_636d.drugtemist-it-fasttext-75-nerhttps://temu.bsc.es/multicardioner/drugtemist-en-fasttext-85-nerhttps://temu.bsc.es/multicardioner/drugtemist-en-fasttext-8-nerhttps://temu.bsc.es/multicardioner/wi_generate_fasttext_traininge1_code_fasttext_qwqdistemist-fasttext-9-nerhttps://temu.bsc.es/multicardioner/symptemist-fasttext-8-nerhttps://temu.bsc.es/symptemist/gru_fasttext_model
Gojek Statement Review Classifier
This application is designed to classify review statements into positive, neutral, or negative sentiments using traditional machine learning and deep learning models, built on Gojek review data
See more on web demo, and github:
[1] https://gojek-sentiment-review-classifier-kelompok6.streamlit.app/
[2] https://github.com/FebryantoAdityaRizky020204/gojek-sentiment-review-classifier/tree/main
FastTextCatopenlid-45-fasttextinstruction_filtering_fasttext_per_domain_seed_data_math_w_openthoughtsb2_science_fasttext_neg_wikipedia
