datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hnet-chunking-results
H-Net dynamic-chunking: experiment results
Every result table behind the study, including the ones that failed. 41 experiment directories;
each has a RESULTS.md (verdict + caveats) alongside its machine-readable CSV/JSON, and the
82.5M re-runs also carry an auto-rendered RESULTS_auto.md produced by the same report script
as the pilot tables.
Headline findings
experiment
question
verdict
exp21_tier1_pilot
does a parity objective equalise chunk… See the full description on the dataset page: https://huggingface.co/datasets/AdaptiveChunking/hnet-chunking-results.Llama-3.2-1B-beam-search_4_no_chunking_H200Llama-3.2-1B-bon_256_no_chunking_H200Falcon3-1B-bon_64_no_chunking_H200Llama-3.2-1B-bon_64_no_chunking_H200Falcon3-1B-bon_256_no_chunking_H200Llama-3.2-1B-beam-search_16_no_chunking_H200Falcon3-1B-beam-search_16_no_chunking_H200Llama-3.2-1B-dvts_4_no_chunking_H200evalap-mediatechs-legi-chunking-evaluation-v1-116
MediaTech's LEGI Chunking Evaluation V1 (ID: 116)
Evaluation of severals chunking strategies for MediaTech's LEGI dataset.
Overview
This dataset contains 51 experiments
from the EvalAP evaluation platform.
Datasets: LEGI Synthetic QA Dataset
Metrics: contextual_precision, contextual_recall, contextual_relevancy, faithfulness, judge_precision
Scores
LEGI Synthetic QA Dataset
model
contextual_precision
contextual_recall
contextual_relevancy… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mediatechs-legi-chunking-evaluation-v1-116.Falcon3-1B-beam-search_4_no_chunking_H200Llama-3.2-1B-dvts_256_no_chunking_H200Llama-3.2-1B-beam-search_64_no_chunking_H200Falcon3-1B-dvts-4_no_chunking_H200Falcon3-1B-dvts-16_no_chunking_H200Llama-3.2-1B-dvts_64_no_chunking_H200Falcon3-1B-bon_16_no_chunking_H200Llama-3.2-1B-bon_4_no_chunking_H200Llama-3.2-1B-bon_16_no_chunking_H200Llama-3.2-1B-dvts_16_no_chunking_H200Falcon3-1B-beam-search_64_no_chunking_H200evalap-mediatechs-service-public-travail-emploi-chunking-evaluation-v1-117
MediaTech's Service Public & Travail Emploi Chunking Evaluation V1 (ID: 117)
Evaluation of severals chunking strategies for MediaTech's Service Public & Travail Emploi datasets.
Overview
This dataset contains 24 experiments
from the EvalAP evaluation platform.
Datasets: MFS_good_dataset
Metrics: contextual_precision, contextual_recall, contextual_relevancy, faithfulness, judge_notator
Scores
MFS_good_dataset
model
contextual_precision… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mediatechs-service-public-travail-emploi-chunking-evaluation-v1-117.Falcon3-1B-dvts-256_no_chunking_H200Falcon3-1B-bon_4_no_chunking_H200repoexec_chunking_bm25etri-chunkingFalcon3-1B-dvts-64_no_chunking_H200evalap-mediatechs-service-public-travail-emploi-chunking-evaluation-v1-117
MediaTech's Service Public & Travail Emploi Chunking Evaluation V1 (ID: 117)
Evaluation of severals chunking strategies for MediaTech's Service Public & Travail Emploi datasets.
Overview
This dataset contains 24 experiments
from the EvalAP evaluation platform.
Datasets: MFS_good_dataset
Metrics: contextual_precision, contextual_recall, contextual_relevancy, faithfulness, judge_notator
Scores
MFS_good_dataset
model
contextual_precision… See the full description on the dataset page: https://huggingface.co/datasets/kaaloo/evalap-mediatechs-service-public-travail-emploi-chunking-evaluation-v1-117.etri-chunking-cleaned
