CLM
Datasets
All datasets matching “CLM”CLM-v0.1-Pretrain-Nemotronclmm-benchmarkdeepswe-clm-train-embeddings-8k
DeepSWE PRM training embeddings (Qwen3-8B, 8k)
Frozen Qwen3-8B last-token-pooled embeddings (4096-d, float16) of every step of the
DeepSWE training-pool rollouts: 405,919 steps · 4,701 trajectories · 113 tasks.
This is the data the released DeepSWE PRM heads
(tarsur385/deepswe-prm-heads-8k) were fine-tuned on.
Embedded with preprocessing/deepswe/embed_shard.py at max_model_len 8192: the state is the
chat-templated step context truncated to its last 8191 tokens; the action is the… See the full description on the dataset page: https://huggingface.co/datasets/Contrastive-LM/deepswe-clm-train-embeddings-8k.cl-maven
MAVEN — continual learning splits
The MAVEN corpus cut into the task streams used by the OpenED continual-learning
experiments (CED / continual event detection, 5 tasks per permutation, 5 permutations).
Layout
raw/{train,dev,test}.jsonl the corpus before it is cut into tasks
perm<k>/streams.json the label groups of permutation k, in task order
perm<k>/<task>/{train,dev,test}.jsonl
train.jsonl of task t holds that task's labels plus a replay… See the full description on the dataset page: https://huggingface.co/datasets/datht/cl-maven.romansh-grischun-morphological-corpus
Romansh Grischun Morphological Corpus
A morphologically annotated corpus of Rumantsch Grischun, the standardized written variety of Romansh.
Dataset Summary
This dataset provides gold-standard morphosyntactically annotated and lemmatized corpora for Rumantsch Grischun (standard written Romansh, ISO 639-3: roh), a national language of Switzerland.
The primary annotation layer preserves the rich Xerox/Foma two-level morphological tags used by the finite-state… See the full description on the dataset page: https://huggingface.co/datasets/simon-clmtd/romansh-grischun-morphological-corpus.clmet_3_1
Dataset Card for clmet_3_1
NOTES:
Some of the annotations in the class and pos configs are not properly formed. These are indicated with warning messages when the dataset is loaded.
In addition to the classes mentioned in the README for the dataset, there is an additional class in the class dataset called QUOT. As far as I can tell, this is used for tagging all quotation marks
When the class and pos configs are loaded, the available class/pos tags are shown at the top… See the full description on the dataset page: https://huggingface.co/datasets/biglam/clmet_3_1.
