datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nerf-synthetic-mirrornerf-gs-datasetsI keep a collection compiled of existing datasets from various sources for training NeRFs or Splats. This dataset is most of that collection. All of the individual scenes also have a trained Gaussian Splat.
https://rishit-dagli.github.io/2025/03/28/nerf-gs-datasets.html
ner-jsonlnerfew-nerd
Dataset Card for "Few-NERD"
#dataset-description)
Dataset Summary
Supported Tasks and Leaderboards
Languages
Dataset Structure
Data Instances
Data Fields
Data Splits
Dataset Creation
Curation Rationale
Source Data
Annotations
Personal and Sensitive InformationConsiderations for Using the Data
Social Impact of Dataset
Discussion of Biases
Other Known Limitations
Additional Information
Dataset Curators
Licensing Information
Citation Information
Contributions
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DFKI-SLT/few-nerd.bad_prompt
Negative Embedding / Textual Inversion
Idea
The idea behind this embedding was to somehow train the negative prompt as an embedding, thus unifying the basis of the negative prompt into one word or embedding.
Side note: Embedding has proven to be very helpful for the generation of hands! :)
Usage
To use this embedding you have to download the file aswell as drop it into the "\stable-diffusion-webui\embeddings" folder.
Please put the embedding in… See the full description on the dataset page: https://huggingface.co/datasets/Nerfgun3/bad_prompt.kpwr-nerKPWR-NER tagging dataset.gutenberg_spacy-ner
Dataset Card for "gutenberg_spacy-ner"
More Information needed
Mip-NeRF360VibeVoicesample-ner
Dataset Card for "sample-ner"
More Information needed
open-ner-standardized
Dataset Card for OpenNER 1.0
OpenNER 1.0 is a standardized collection of openly-available named entity recognition (NER) datasets.
OpenNER contains 36 NER corpora that span 52 languages, human-annotated in varying named entity ontologies.
We correct annotation format issues, standardize the original datasets into a uniform representation with consistent entity type names across corpora, and provide the collection in a structure that enables research in multilingual and… See the full description on the dataset page: https://huggingface.co/datasets/bltlab/open-ner-standardized.nerfbaselines-dataProfNER_corpus_NER
Description
Gold standard annotations for profession detection in Spanish COVID-19 tweets
The entire corpus contains 10,000 annotated tweets. It has been split into training, validation, and test (60-20-20). The current version contains the training and development set of the shared task with Gold Standard annotations. In addition, it contains the unannotated test, and background sets will be released.
For Named Entity Recognition, profession detection, annotations are distributed… See the full description on the dataset page: https://huggingface.co/datasets/Biomedical-TeMU/ProfNER_corpus_NER.nerf-syntheticdatasetssimplevla-grpo-assets
SimpleVLA GRPO Grasp Assets
This dataset contains the released USD object assets used by the SimpleVLA-style GRPO grasping experiments.
Expected local layout after running scripts/download_assets.sh:
/data4/nerako/reasoning/RLinf_assets/grasp_assets/
nerf-syntheticnerf_syntheticnerfbaselines-supplementaryNeRSemblecantemist-nerhttps://temu.bsc.es/cantemist/tlunified-ner
🪐 spaCy Project: TLUnified-NER Corpus
Homepage: Github
Repository: Github
Point of Contact: ljvmiranda@gmail.com
Dataset Summary
This dataset contains the annotated TLUnified corpora from Cruz and Cheng
(2021). It is a curated sample of around 7,000 documents for the
named entity recognition (NER) task. The majority of the corpus are news
reports in Tagalog, resembling the domain of the original ConLL 2003. There
are three entity types: Person (PER), Organization… See the full description on the dataset page: https://huggingface.co/datasets/ljvmiranda921/tlunified-ner.gutenberg_spacy-nerrobust-e-nerf
Robust e-NeRF Synthetic Event Dataset
This repository contains the synthetic event dataset used in Robust e-NeRF to study the collective effect of camera speed profile, contrast threshold variation and refractory period on the quality of NeRF reconstruction from a moving event camera. The dataset is simulated using an improved version of ESIM with three different camera configurations of increasing difficulty levels (i.e. easy, medium and hard)… See the full description on the dataset page: https://huggingface.co/datasets/wengflow/robust-e-nerf.cross_nerCrossNER is a fully-labeled collected of named entity recognition (NER) data spanning over five diverse domains
(Politics, Natural Science, Music, Literature, and Artificial Intelligence) with specialized entity categories for
different domains. Additionally, CrossNER also includes unlabeled domain-related corpora for the corresponding five
domains.
For details, see the paper:
[CrossNER: Evaluating Cross-Domain Named Entity Recognition](https://arxiv.org/abs/2012.04373)Pile-NER-type
Intro
Pile-NER-type is a set of GPT-generated data for named entity recognition using the type-based data construction prompt. It was collected by prompting gpt-3.5-turbo-0301 and augmented by negative sampling. Check our project page for more information.
License
Attribution-NonCommercial 4.0 International
swedish_ner_corpusWebbnyheter 2012 from Spraakbanken, semi-manually annotated and adapted for CoreNLP Swedish NER. Semi-manually defined in this case as: Bootstrapped from Swedish Gazetters then manually correcte/reviewed by two independent native speaking swedish annotators. No annotator agreement calculated.Voxpopuli_NER
VoxPopuli_NER
VoxPopuli-NER is derived from the VoxPopuli corpus and specifically enhanced for
Named Entity Recognition (NER) tasks focusing on political and geographical entities.
It includes 879 audio samples, annotated with 2469 unique entity types. The dataset consists of the English part of the test set of VoxPopuli.
See full details in the WhisperNER paper.
citation
If you find this usful, please cite the following works:
@article{ayache2024whisperner… See the full description on the dataset page: https://huggingface.co/datasets/aiola/Voxpopuli_NER.deblur-e-nerf
Deblur e-NeRF Synthetic Event Dataset
This repository contains the synthetic event dataset used in Deblur e-NeRF to study the collective effect of camera speed and scene illuminance on the quality of NeRF reconstruction from a moving event camera. It is an extension of the synthetic event dataset used in Robust e-NeRF. The dataset is simulated using an improved version of ESIM with three different camera configurations of increasing difficulty… See the full description on the dataset page: https://huggingface.co/datasets/wengflow/deblur-e-nerf.
