consistency
ConsistencyAnnotationseed_code_self_consistencyFIRM-Edit-Consistencyknowledge_consistency_of_LLMs
What it is:
Each dataset in this delivery is made up of query clusters that test an aspect of the consistency of the LLM knowledge about a particular domain. All the questions in each
cluster are meant to be answered either 'yes' or 'no'. When the answers vary within a cluster, the knowledge is said to be inconsistent. When all the questions in a cluster
are answered 'no' when the expected answer is 'yes' (or viceversa), the knowledge is said to be 'incomplete' (i.e., maybe the LLM… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/knowledge_consistency_of_LLMs.question-consistency-datasets
Question Consistency — concept datasets
Item pools for the question-consistency
preference/judgement-elicitation harness (forced-choice pairwise comparisons → Thurstonian fit →
consistency metrics). Each config is a flat list of items in a single item (string) column.
config
rows
what
items_500
500
500-concept sentiment/judgement pool
items_2000
2000
2000-concept pool (large-scale runs)
curated_concepts
250
curated rich multi-word concepts spanning categories… See the full description on the dataset page: https://huggingface.co/datasets/arcadia-impact/question-consistency-datasets.task066_timetravel_binary_consistency_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task066_timetravel_binary_consistency_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task066_timetravel_binary_consistency_classification.
