Team Ai
20 results

language models

prompt-agnostic-language-models /pal-results0 likes579 downloads3mo agoHugging FaceCCB /cis5300-language-models CIS 5300 Language Models Dataset Dataset for Homework 3 of CIS 5300 (Natural Language Processing) at Penn. Cities config Country-of-origin classification over short city-name strings, drawn from nine countries (Afghanistan, China, Germany, Finland, France, India, Iran, Pakistan, South Africa). from datasets import load_dataset cities = load_dataset("CCB/cis5300-language-models", "cities") Split Rows Has labels? train 12,392 yes validation 1,548 yes test 1… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-language-models.text10K<n<100K0 likes475 downloads5mo agoHugging Facebeatsprom /multimodal-vision-language-video-models-2026 👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition) A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators. Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.tabularfeature-extractionn<1K8 likes169 downloads2mo agoHugging FaceSetloop /What-Gradients-Add-to-Text-Leakage-in-Split-Language-Models-Counted-per-Token-and-per-Documentgated Paper A: reproduction and peer-review release This dataset contains the final paper, the evidence used for its three experiments, and the historical material explicitly discussed in the paper. Final paper The current branch includes the 2 October 2026 punctuation revision of the 26-page manuscript. Its scientific content and experimental evidence are unchanged. Tag v1.0 retains the frozen release and its original manuscript; the updated PDF and source package are… See the full description on the dataset page: https://huggingface.co/datasets/Setloop/What-Gradients-Add-to-Text-Leakage-in-Split-Language-Models-Counted-per-Token-and-per-Document.3 likes167 downloads8d agoHugging Facehassan-wajid /Spatial-Blind-Spots-in-Vision-Language-Modelslicense: mit model_evaluated: name: Qwen3-VL-2B-Instruct url: https://huggingface.co/Qwen/Qwen3-VL-2B-Instruct evaluation_notebook: https://www.kaggle.com/code/wajidhassanmoosa/blind-spot-qwen3-2b evaluation_setup: | The model evaluated in this study is Qwen3-VL-2B-Instruct. Evaluation was conducted using the Hugging Face Transformers library with automatic device mapping (device_map="auto") and "bfloat16" dtype selection. For each example: The image was provided as part of a… See the full description on the dataset page: https://huggingface.co/datasets/hassan-wajid/Spatial-Blind-Spots-in-Vision-Language-Models.imagen<1K7 likes148 downloads7mo agoHugging Faceastha /languagemodelsforRNNdecompositionThis repository is for the paper "Decomposing a Recurrent Neural Network into Modules for Enabling Reusability and Replacement". To use the data, there are two directories: language datasets: Contains the necessary Tatoeba files used for the experiments. We have experimented with 4 languages(English, French, Italian and German). language_models: Contains all trained language models and scripts to train them. It's organized in this way: language_models/{X}: contains language models for X… See the full description on the dataset page: https://huggingface.co/datasets/astha/languagemodelsforRNNdecomposition.text100K<n<1M1 likes115 downloads4y agoHugging Face