chunking
distilbert-cross-segment-document-chunkingfewshot-learning-bart-base-paraphrase-finetuned-for-chunkingChunkingChipShotstext-seg-lm-qwen2-0.5b-cot-topic-chunkingSentence-Chunking-Afri_BERTA_amharic_longtexttext-seg-lm-qwen2-0.5b-summary-chunkingqwen2_chunking_mlp_freeze_uniform_with_shared_start_and_end_2_12_ptqwen2_chunking_mlp_freeze_uniform_with_shared_start_and_end_2_12_sft
Datasets
All datasets matching “chunking”hnet-chunking-probes
H-Net chunker boundary probes
Longitudinal boundary decisions for 31 H-Net runs, logged on a fixed, byte-identical
FLORES+ probe at every checkpoint. This is the raw material for studying when a learned
segmentation stabilises.
Layout
<run>/{step:06d}__{lang}.npz, plus <run>/probe_text.jsonl (the raw probe text, so byte
offsets can be aligned to external gold data).
40 log-spaced steps: 0, 1, 2, 4, 8, 16, 32, 64, 128, 200, then every 200 to 6000.
The early… See the full description on the dataset page: https://huggingface.co/datasets/AdaptiveChunking/hnet-chunking-probes.act-chunking-study
ACT action chunking under delay and disturbance: per-episode results
This repository holds the raw outputs of the study
ACT action chunking under delay and disturbance.
Every simulated episode behind the study's tables is stored as one JSON record. The repository also
holds the scripts that the Hugging Face Jobs ran.
code/scripts/ holds the scripts as uploaded for the jobs. The GitHub repository is the maintained copy.
results/<run>/meta.json holds the settings of a run: grid… See the full description on the dataset page: https://huggingface.co/datasets/FTG64/act-chunking-study.chunking-test-datahnet-chunking-results
H-Net dynamic-chunking: experiment results
Every result table behind the study, including the ones that failed. 41 experiment directories;
each has a RESULTS.md (verdict + caveats) alongside its machine-readable CSV/JSON, and the
82.5M re-runs also carry an auto-rendered RESULTS_auto.md produced by the same report script
as the pilot tables.
Headline findings
experiment
question
verdict
exp21_tier1_pilot
does a parity objective equalise chunk… See the full description on the dataset page: https://huggingface.co/datasets/AdaptiveChunking/hnet-chunking-results.rag-chunking-ablation-demo-assets
RAG chunking demo assets (Natural Questions bench + prebuilt FAISS indices)
Prebuilt, read-only assets that let an interactive retrieval demo start
without embedding a corpus at boot.
Contents
path
what it is
data/nq/large_n1000/docs.jsonl
1000 Natural Questions Wikipedia documents, pre-split into sentences (id, title, sentences)
data/nq/large_n1000/questions.jsonl
1032 NQ validation questions with short answers and gold document titles… See the full description on the dataset page: https://huggingface.co/datasets/sfczaa/rag-chunking-ablation-demo-assets.Beyond-Chunking
Beyond Chunking: Node Integrity and Field Structure in Bengali Retrieval-Augmented Agent Memory
Official anonymized reproduction suite for the paper:
"Beyond Chunking: Node Integrity and Field Structure in Bengali Retrieval-Augmented Agent Memory"
Overview
This repository provides complete, self-contained artifacts, datasets, experimental code, and evaluation oracles to independently verify all empirical claims, tables, and mechanistic analyses presented in… See the full description on the dataset page: https://huggingface.co/datasets/RaiyanKhaan/Beyond-Chunking.
