compact
Datasets
All datasets matching “compact”compact-alignments
compact-alignments — per-verse, per-book, content-addressed
The token-position companion to lexeme-alignments (which is
aggregated/type-level and can't tell you what happened in any one verse). This dataset restores
position: for a given edition's Bible book, which Hebrew/Greek content word aligned to which
target-text token, verse by verse.
The authoritative list of what's published is always manifest.json, not this file.
Original-language source editions (needed… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/compact-alignments.slm-parameter-audit
SLM card-vs-artifact parameter audit
An autonomous audit of small-language-model repos on the Hugging Face Hub. For each
in-scope model (independent builders training very small models from scratch, roughly
0.5M–500M parameters), the parameter count stated in the model card is compared against
the actual artifact: the safetensors header, config.json, and the training script where
present. A mismatch is recorded when the card's number does not match the artifact's
real parameter… See the full description on the dataset page: https://huggingface.co/datasets/Compactbot/slm-parameter-audit.CompactDS-102GB
Introduction
CompactDS is a diverse, high-quality, web-scale datastore that achieves high retrieval accuracy and subsecond latency on a single-node deployment, making it suitable for academic use. Its core design combines a compact set of high-quality, diverse data sources with in-memory approximate nearest neighbor (ANN) retrieval and on-disk exact search. We release CompactDS and our retrieval pipeline as a fully reproducible alternative to commercial search, supporting future… See the full description on the dataset page: https://huggingface.co/datasets/alrope/CompactDS-102GB.compact-scientific-lm-dataFlexiSLM-Data-2M-s2s-compact
FlexiSLM-Data — Speech-to-Speech Part (2.43M filtered samples, 385G in size)
Paper: https://arxiv.org/abs/2606.31247
Demo page: https://flexislm.github.io/
Code: https://github.com/AmphionTeam/FlexiSLM
FlexiSLM-Data is a large-scale, single-turn English speech-to-speech dialogue dataset
for training FlexiSLM, a spoken language model.
This repository contains the paired prompt-and-response audio portion of the release in
WebDataset format.
Related data releases… See the full description on the dataset page: https://huggingface.co/datasets/FlexiSLM/FlexiSLM-Data-2M-s2s-compact.agent-memory-compaction-trajectories
Agent Memory Compaction Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/agent-memory-compaction-trajectories.
