Team Ai
20 results

language-model

prompt-agnostic-language-models /pal-results0 likes600 downloads3mo agoHugging FaceCCB /cis5300-language-models CIS 5300 Language Models Dataset Dataset for Homework 3 of CIS 5300 (Natural Language Processing) at Penn. Cities config Country-of-origin classification over short city-name strings, drawn from nine countries (Afghanistan, China, Germany, Finland, France, India, Iran, Pakistan, South Africa). from datasets import load_dataset cities = load_dataset("CCB/cis5300-language-models", "cities") Split Rows Has labels? train 12,392 yes validation 1,548 yes test 1… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-language-models.text10K<n<100K0 likes497 downloads5mo agoHugging FaceSakonii /nepalitext-language-model-dataset Dataset Card for "nepalitext-language-model-dataset" Dataset Summary "NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia. Supported Tasks and Leaderboards This dataset is intended to pre-train language models and word representations on Nepali Language. Languages The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.texttext-generation10M<n<100M8 likes480 downloads1y agoHugging Facekarthik-2905 /core-language-model-data Core Language Model A small coding + general-purpose language model built from scratch in pure NumPy — no PyTorch, no JAX, no autograd library. Every piece (autograd engine, tokenizer, attention, layer norm) is hand-written and gradient-checked. This is a from-first-principles build of a transformer language model, aimed at understanding every formula well enough to reimplement it without reference. Why pure NumPy? The goal isn't a competitive model — it's a fully… See the full description on the dataset page: https://huggingface.co/datasets/karthik-2905/core-language-model-data.0 likes222 downloads3mo agoHugging FaceUnified-Language-Model-Alignment /Anthropic_HH_Golden Dataset Card for Anthropic_HH_Golden This dataset is constructed to test the ULMA technique as mentioned in the paper Unified Language Model Alignment with Demonstration and Point-wise Human Preference (under review, and an arxiv link will be provided soon). They show that replacing the positive samples in a preference dataset by high-quality demonstration data (golden data) greatly improves the performance of various alignment methods (RLHF, DPO, ULMA). In particular, the ULMA… See the full description on the dataset page: https://huggingface.co/datasets/Unified-Language-Model-Alignment/Anthropic_HH_Golden.text10K<n<100K40 likes218 downloads3y agoHugging FacePlim /language_model_frtext1M<n<10M0 likes207 downloads4y agoHugging Face