datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fact-checking-abstention-saes-data
Fact-Checking Abstention SAEs — data
Companion data for lucasfrag/fact-checking-abstention-saes.
Per base model (llama-3.1-8b-instruct/, qwen3-8b/)
file
content
examples.jsonl
the exact ordered list of training claims (VitaminC + FEVER train splits, gold evidence), one JSON per line: claim, evidence (list), label (SUP/REF/NEI), source. Line i is prompt i of the harvest, so activations can be regenerated deterministically.
decisions.safetensors… See the full description on the dataset page: https://huggingface.co/datasets/lucasfrag/fact-checking-abstention-saes-data.Multi_News_fact_checking_claims
Dataset Card for "v2"
More Information needed
task966_ruletaker_fact_checking_based_on_given_context
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task966_ruletaker_fact_checking_based_on_given_context
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task966_ruletaker_fact_checking_based_on_given_context.vietnamese-fact-checking-verifier-data-v3-1
Vietnamese fact-checking verifier data
Leakage-aware document-level 80/10/10 split derived from
aiMy144/vietnamese-fact-checking-claims-v3-1 at revision e20c1eddcbb4a5de862dbc6965eabee43c9766da.
Input is (evidence_text, claim) and labels are SUPPORTED, REFUTED, and
NOT_ENOUGH_INFO. Exact duplicate claims are retained only once.
portuguese-fact-checking
Portuguese Automated Fact-Checking
Fake.BR
COVID19.BR
MuMiN-PT
Info (fake/true)
🖥️
💬
X
Domain
General
Health
"General" (Health)
Year
2016–2018
2020
2020–2022
Approach [1]
bottom-up
bottom-up
top-down
Size
3580/3580
848/1139
1339/65
% URL
1.0%/0.7%
28.9%/56.9%
0.3%/0.0%
Avg. # words
181.4/183.1
167.7/111.1
18.9/16.9
Corpora characteristics after cleaning. Top-down starts with fact-checked claims; bottom-up seeks for new misinformation in posts.… See the full description on the dataset page: https://huggingface.co/datasets/ju-resplande/portuguese-fact-checking.factguard_factchecking_datasetstwitter_factchecking_testvietnamese-fact-checking-verifier-data
Vietnamese fact-checking verifier data
Leakage-aware document-level 80/10/10 split derived from
Loctran123/vietnamese-fact-checking-claims at revision 63963b973af864ecc20313636b1786c31bbb4a41.
Input is (evidence_text, claim) and labels are SUPPORTED, REFUTED, and
NOT_ENOUGH_INFO. Exact duplicate claims are retained only once.
factchecking-calibration-datasetfact-checking-vietnamese-newsviet-fact-checking
Vietnamese Evidence Corpus for Fact-Checking & RAG (v1.0)
This dataset is a clean, standardized, and unified Vietnamese Evidence Corpus (v1.0) built for research in Information Retrieval, Retrieval-Augmented Generation (RAG), and Fact-Checking / Claim Verification.
Dataset Statistics
Total Documents: 13,572 (frozen unique records, duplicates filtered out)
Languages: ~70% Vietnamese (vi), ~30% English (en)
Size: 115.33 MB
Documents by Source… See the full description on the dataset page: https://huggingface.co/datasets/aiMy144/viet-fact-checking.fact-checking-embedding-dataresplit_multi_fact_checking_datasetFactCheckingEvalrlvr_task966_ruletaker_fact_checking_based_on_given_contextvietnamese-fact-checking-claims
Vietnamese Fact-Checking Claims
Generated claim-verification data derived from the Vietnamese Evidence Corpus.
Each article contains claims labeled as supported, refuted, or not having enough
information, together with evidence and a short rationale.
Statistics
12,238 source articles
73,454 generated claims
24,476 SUPPORTED claims
24,502 REFUTED claims
24,476 NOT_ENOUGH_INFO claims
Main fields
Article: id, date_iso, full_text, claims
Claim: claim… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-fact-checking-claims.flan_combined_task966_ruletaker_fact_checking_based_on_given_contextDisaster-Type_Classification_Dataset_for_Automated_Fact-Checking
DTCD-AFC: Disaster-Type Classification Dataset for Automated Fact-Checking
Overview
The DTCD-AFC is a dataset designed for disaster-type classification evaluation for automated fact-checking.
It consists of multimodal social media posts collected based on past natural disasters, each labeled with the disaster type to which its content relates.
The social media posts are sourced from CrisisMMD.
Files
disaster_type_classification_dataset_for_afc.csv: The CSV… See the full description on the dataset page: https://huggingface.co/datasets/o-yas/Disaster-Type_Classification_Dataset_for_Automated_Fact-Checking.qwen3_0.6b-rlvr_task966_ruletaker_fact_checking_based_on_given_contextresplit_multi_fact_checking_dataset_resampling
