datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mu-glioma-post-processed
Processed MU-Glioma-Post
Start Here
Use manifests/experiment_index.csv as the main case-level table.
Use manifests/longitudinal_index.csv when ordering patient timepoints for longitudinal work.
Use manifests/validation_summary.json to confirm the processed tree is complete.
Directory Guide
images_native/
Canonical modality symlinks to the source images.
File names use t1, t1c, flair, and t2.
images_reoriented/
Not populated because audit showed all volumes… See the full description on the dataset page: https://huggingface.co/datasets/sbandred/mu-glioma-post-processed.ave-speech-lipemg-processed
AVE Speech — preprocessed for lip–EMG fusion
Code: diddmstjr07/silent-speech-viseme-emg — model definition, training, evaluation, and the analyses behind these numbers.
Related releases: checkpoints · AVE preprocessed · Confusable-100
A derivative of the AVE Speech corpus
(Zhou et al., IEEE THMS 2025), preprocessed into the exact form used to train the
lip–EMG fusion models in the companion work. This is not new recorded data; it is the
original corpus with the preprocessing… See the full description on the dataset page: https://huggingface.co/datasets/diddmstjr/ave-speech-lipemg-processed.processed_test_splitsqelathrym-hsin-operational-fractal-process-geometry
QELATHRYM–HSIN: Operational Fractal Process Geometry
Task-dependent fractal geometry · Recursive operator trees · Exact task retention · Finite neural tangent representations
Research version 5.0.0 · 1 October 2026 · Self-contained, unreviewed mathematical research draft
QELATHRYM–HSIN proposes a geometry of which quadratic tasks remain accessible through recursive branching under a specified local readout rule. The research connects normalized operator trees, inverse-adjoint… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/qelathrym-hsin-operational-fractal-process-geometry.3d-pinn-dim55-processed_datasetnvidia_openmathinstruct-2-simple-processed元データ
https://huggingface.co/datasets/nvidia/OpenMathInstruct-2
repro-ai4slt-empirical-processes-in-lean-4-for-formal-statistical-learning-theory-traces
Agent traces
Agent sessions published from a Trackio Logbook.
youtube_audio_processed_dataset
YouTube Audio Processed Dataset
This dataset contains high-quality segmented speech datasets preprocessed from YouTube videos using the Emilia preprocessor framework.
Dataset Structure
Each row in the dataset contains a segmented audio clip, its aligned high-fidelity transcript, speaker labeling, and objective speech quality assessment scores.
Features
file_name: Audio column containing the relative path to the segmented .mp3 clip.
text: The… See the full description on the dataset page: https://huggingface.co/datasets/Aarjanm/youtube_audio_processed_dataset.bigclonebench-processedrepro-tighter-regret-lower-bound-for-gaussian-process-bandits-with-squared-exponential-traces
Agent traces
Agent sessions published from a Trackio Logbook.
gold-silver-mineral-process-cpt-candidates
Gold/silver mineral-process CPT candidates
English raw documents (text) about gold/silver and transferable hard-rock mineral processing.
Source: BAAI/IndustryCorpus2_mining revision bf358a2f8105e4ac468141796e5a1a530685ae2e. English only.
These are documents, not chat pairs.
Configs
Config
Rows
Notes
default
59,749
all English bands
english_high
18,864
publisher quality 4.00–4.59
english_middle
34,601
publisher quality 3.00–4.00
english_low
6,284… See the full description on the dataset page: https://huggingface.co/datasets/hicham-taoufik/gold-silver-mineral-process-cpt-candidates.process-revision-sci-writewarehouse-process-mining-benchmark
Warehouse Process Mining Benchmark
Benchmark results for WareFlowTwin process mining and optimization, aligned with VillanovaAI/Temporal_Logistics_Inventory_Movements.
Contents
File
Description
manifest.json
Dataset metadata
eval_results.json
Aggregate benchmark metrics
warehouse_*.json
Per-instance process mining + optimization results
Pipeline Evaluated
Event log construction (lot_id, case_id, activity, timestamp, warehouse… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/warehouse-process-mining-benchmark.research-papers-dataset-mixtral7B-processed2
Research Papers Dataset - Processed with Train/Test/Valid Splits
This dataset contains preprocessed research papers with the following enhancements, split into train/test/validation sets.
Dataset Splits:
Train: 7,328 entries (85.0%)
Test: 431 entries (5.0%)
Valid: 863 entries (10.0%)
Preprocessing Applied:
Section Splitting: Papers are split into logical sections (Abstract, Introduction, Methods, Results, etc.)
Whitespace Normalization: Excessive whitespace… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/research-papers-dataset-mixtral7B-processed2.trivia-qa-kg-processedProcess_datarecursion-process
CatQualia recursion process corpus — self-improvement loop traces
4,631 rows · 5,750,770 bytes · JSON Lines, one object per line.
What this is
Traces of a self-improvement loop in operation: what the loop proposed, what the measurement returned, and what was kept. Directly relevant to the self-falsification subject matter of the wider corpus.
Schema
Fields of the first record, read from the file in this repository:
Field
Type
cycle_id
str… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/recursion-process.processed2hotpot-qa-kg-processedNovaciano__BLAST_PROCESSING-3.2-1B-details
Dataset Card for Evaluation run of Novaciano/BLAST_PROCESSING-3.2-1B
Dataset automatically created during the evaluation run of model Novaciano/BLAST_PROCESSING-3.2-1B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Novaciano__BLAST_PROCESSING-3.2-1B-details.musique-kg-processedagentica-org_deepscaler-preview-dataset-simple-processed元データセット
https://huggingface.co/datasets/agentica-org/DeepScaleR-Preview-Dataset
tw-processed-related-law-article
Dataset Card for tw-processed-related-law-article
本資料集將中華民國(臺灣)現行法規以「條」為單位重新切分,並把每一條與其引用之相關條文一起合併呈現,方便用於法律語料的持續預訓練(continued pretraining)與條文檢索/問答(retrieval / QA)任務。
Dataset Details
Dataset Description
資料來自中華民國公開法規與條文內容,經過下列處理:
將每部法規依條文(flno)切分為獨立樣本。
從條文內文中解析出「相關法條」,將被引用之條文內容串接於原條文後,形成可獨立閱讀的訓練樣本。
標註每條的母法名稱(name)與法規代碼(pcode)。
最終樣本以「法規名稱:xxx 第 N 條 + 條文 + 相關法條」格式呈現,每一筆都是封閉的條文上下文,能直接做為語言模型訓練的輸入。
Curated by: Huang Liang Hsun
Language(s) (NLP): Traditional… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-processed-related-law-article.research-papers-dataset-mixtral7B-processed
Research Papers Dataset - Processed
This dataset contains preprocessed research papers with the following enhancements:
Preprocessing Applied:
Section Splitting: Papers are split into logical sections (Abstract, Introduction, Methods, Results, etc.)
Whitespace Normalization: Excessive whitespace removed and normalized
Punctuation Fixing: Missing spaces after punctuation marks corrected
Sentence Boundary Fixing: Proper sentence boundaries established
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/research-papers-dataset-mixtral7B-processed.processed_dataset_4hadith-gnn-processed-datamaxwell-jia_aime_2024-simple-processed
