context
Datasets
All datasets matching “context”imagescertificatessql-create-context
Overview
This dataset builds from WikiSQL and Spider.
There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/sql-create-context.barbet-long-context-sft
Barbet long-context SFT
Release ee31706b7783bc2866f8e152d76dfde225effc2029aff24736bbe12cf024205a preserves 4422 active records. This is one joint
assistant-only SFT dataset; no Barbet model training has been run.
The skill-prefill migration has revised 1245
of 1254 records from its fixed base snapshot.
Revisions replace their original records in the explicit shard lists above. Old
bundles and releases remain available at their pinned commits. Additional records
from other… See the full description on the dataset page: https://huggingface.co/datasets/OpenFormosa/barbet-long-context-sft.context_qa_sum_qwen3_synthetic
Context-based QA and Summarization Synthetic Dataset
Overview
This dataset contains synthetic context-based question-answering (QA) and summarization data. The data was synthesized using:
Source context: openbmb/Ultra-FineWeb
Synthesis model: Qwen3-30B-A3B-Instruct-2507
Each context is obtained by taking the initial segment of raw pretraining text from Ultra-FineWeb, truncated to at most the corresponding number of tokens, while ensuring the truncation does not occur in… See the full description on the dataset page: https://huggingface.co/datasets/yuyijiong/context_qa_sum_qwen3_synthetic.paracrawl_context
Dataset Card for ParaCrawl_Context
This is a dataset for document-level machine translation introduced in the ACL 2024 paper Document-Level Machine Translation with Large-Scale Public Parallel Data. It is a dataset consisting of parallel sentence pairs from the ParaCrawl dataset along with corresponding preceding context extracted from the webpages the sentences were crawled from.
Dataset Details
Dataset Description
This dataset adds document-level… See the full description on the dataset page: https://huggingface.co/datasets/Proyag/paracrawl_context.
