datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
law-stack-exchange
Dataset Card for Law Stack Exchange Dataset
Dataset Summary
Dataset from the Law Stack Exchange, as used in "Parameter-Efficient Legal Domain Adaptation".
Citation Information
@inproceedings{li-etal-2022-parameter,
title = "Parameter-Efficient Legal Domain Adaptation",
author = "Li, Jonathan and
Bhambhoria, Rohan and
Zhu, Xiaodan",
booktitle = "Proceedings of the Natural Legal Language Processing Workshop 2022",
month = dec… See the full description on the dataset page: https://huggingface.co/datasets/jonathanli/law-stack-exchange.stackexchangestrain-law-stackexchange
Law Stack Exchange — Training, unified schema
A seeded sample of ymoslem/Law-StackExchange, made into retrieval training pairs and reshaped into the strict schema shared by every dataset in this collection. One of the 15 domain sources (code, medical, science, finance, legal) added to the collection's general sources.
Source
ymoslem/Law-StackExchange @ ab2dbaad9a71
Task
legal question → answer
Domain · languages
legal · eng
Queries / documents / qrels
24,326 /… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/train-law-stackexchange.Law-StackExchange
Law-StackExchange Dataset Details
All StackExchange legal questions and their answers from the Law site, up to 14 August 2023.
The repository includes a notebook for the process using the official StackExchange API.
Citation
@misc{Moslem2023-LawStackExchangeDataset,
author = {Moslem, Yasmin},
title = {Law-StackExchange Dataset},
year = 2023,
url = {https://huggingface.co/datasets/ymoslem/Law-StackExchange},
doi =… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/Law-StackExchange.stack-exchange-paired-scoreMedicalSciences-StackExchangeAll StackExchange questions and their answers from the Medical Sciences site, up to 14 August 2023. The repository includes a notebook for the process using the official StackExchange API.
Stack-Exchange-Aprilstack-exchange-dataset
Overview
This dataset consists of three TSV files, namely: cs.tsv, ds.tsv, and p.tsv.
Each file includes the data for the questions asked on a Stack Exchange (SE) question-answering community, from the creation of the community until May 2021.
cs.tsv --> Computer Science SE
ds.csv --> Data Science SE
p.csv --> Political Science SE
File Structure
Each file has the following columns:
id: the question id
title: the title of the question
body: the body or text of the… See the full description on the dataset page: https://huggingface.co/datasets/habedi/stack-exchange-dataset.stackexchange-space-qa
Stack Exchange Space Q&A
Credit: NASA/DOE/Fermi LAT Collaboration
Part of a dataset collection on Hugging Face.
Dataset description
This dataset is a clean, tabular Q&A corpus of space and astronomy knowledge, derived from two Stack Exchange community Q&A sites: Astronomy Stack Exchange (astronomy.stackexchange.com) and Space Exploration Stack Exchange (space.stackexchange.com). Each row is one question paired with its best answer — either the question's… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/stackexchange-space-qa.japanese-stackexchange
japanese-stackexchange
英語による日本語に関する質問ができる Japanese Stack Exchange のデータダンプ をもとにデータを加工し、質問文と回答文のペアになるように調整した QA データセット。
日本語翻訳された StackExchange ではないです。
データ構造
投稿本文は html2text を使ってマークダウン化されています。その際、
コードブロックは ``` で囲まれるように変更されています。
画像 URL に base64 エンコードされた画像が含まれる場合、 [unk] に置き換えています。
default サブセット
id: 質問投稿の ID
question: 質問投稿
answers: 質問に対する回答投稿のリスト
accepted_answer_id: 質問者に選ばれた回答のID。null の可能性がある
popular_answer_id: もっともスコアが高かった回答のID。null の可能性がある
simple サブセット… See the full description on the dataset page: https://huggingface.co/datasets/p1atdev/japanese-stackexchange.stackexchange-preference-dataa1_science_stackexchange_physics_1k_eval_636d
mlfoundations-dev/a1_science_stackexchange_physics_1k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
17.3
53.0
74.6
0.5
36.5
28.3
10.4
2.6
4.9
AIME24
Average Accuracy: 17.33% ± 1.40%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
10.00%
3
30
2
20.00%
6
30
3
13.33%
4
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/a1_science_stackexchange_physics_1k_eval_636d.law_stackexchange
Dataset Card for "law_stackexchange"
More Information needed
stack_exchange_math_bencha1_science_stackexchange_physics_10k_eval_636d
mlfoundations-dev/a1_science_stackexchange_physics_10k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
16.7
57.7
73.0
0.3
40.0
46.3
13.1
4.7
6.5
AIME24
Average Accuracy: 16.67% ± 1.33%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
23.33%
7
30
2
10.00%
3
30
3
16.67%
5
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/a1_science_stackexchange_physics_10k_eval_636d.Law-StackExchange-Flata1_science_stackexchange_physics_3k_eval_636d
mlfoundations-dev/a1_science_stackexchange_physics_3k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
16.0
57.0
73.4
0.3
40.7
40.4
11.7
3.9
6.1
AIME24
Average Accuracy: 16.00% ± 1.40%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
13.33%
4
30
2
16.67%
5
30
3
13.33%
4
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/a1_science_stackexchange_physics_3k_eval_636d.Law-StackExchange-Deduplicateda1_science_stackexchange_physics_0.3k_eval_636d
mlfoundations-dev/a1_science_stackexchange_physics_0.3k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
13.7
53.0
70.6
0.4
38.5
35.4
7.5
4.3
6.1
AIME24
Average Accuracy: 13.67% ± 1.20%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
20.00%
6
30
2
13.33%
4
30
3
13.33%
4
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/a1_science_stackexchange_physics_0.3k_eval_636d.stack_exchange_multiclass_max_500a1_science_stackexchange_physics_1744691414_eval_1331
mlfoundations-dev/a1_science_stackexchange_physics_1744691414_eval_1331
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
GPQADiamond
JEEBench
MMLUPro
LiveCodeBench
CodeElo
Accuracy
17.7
55.8
75.2
44.3
42.4
31.8
24.0
4.9
AIME24
Average Accuracy: 17.67% ± 1.64%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
13.33%
4
30
2
16.67%
5
30
3
10.00%
3
30
4… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/a1_science_stackexchange_physics_1744691414_eval_1331.a1_science_stackexchange_physics_eval_636d
mlfoundations-dev/a1_science_stackexchange_physics_1744737192_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
19.7
57.5
75.6
31.0
43.9
42.6
22.7
5.3
7.6
AIME24
Average Accuracy: 19.67% ± 0.57%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
20.00%
6
30
2
20.00%
6
30
3… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/a1_science_stackexchange_physics_eval_636d.stack_exchange_multilabel_max_500b2_science_embedding_stackexchange_physicsa1_science_stackexchange_physics_3k_eval_2e29cleaned_stack_exchange_python_evalb2_science_length_filtering_stackexchange_physicsa1_science_stackexchange_physics_0.3k_eval_2e29Data-StackExchange-Dedupedstack-exchange-questions-scraper-sample-data
Stack Exchange Questions Scraper
Scrape questions from Stack Overflow and any Stack Exchange site by tag: title, tags, score, views, answers, author and dates. Monitor questions about your product, tech or competitors on a schedule.
What the actor scrapes
❓ Stack Exchange Questions Scraper — Stack Overflow Q&A Data by Tag to JSON & CSV Scrape questions from Stack Overflow and any Stack Exchange site using the official Stack Exchange API. This Stack Overflow… See the full description on the dataset page: https://huggingface.co/datasets/logiover/stack-exchange-questions-scraper-sample-data.
