Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BramVanroy /fineweb-2-duckdbs DuckDB datasets for (dump, id) querying on FineWeb 2 This repo contains some DuckDB databases to check whether a given WARC UID exists in a FineWeb-2 dump. Usage example is given below, but note especially that if you are using URNs (likely, if you are working with CommonCrawl data), then you first have to extract the UID (the id column is of type UUID in the databases). Download All files: huggingface-cli download BramVanroy/fineweb-2-duckdbs --local-dir… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/fineweb-2-duckdbs.0 likes24k downloads1y agoHugging Face02BramVanroy /fineweb-duckdbs DuckDB datasets for id querying on FineWeb This repo contains DuckDB databases to check whether a given WARC UID exists in a FineWeb dump. Usage example is given below, but note especially that if you are using URNs (likely, if you are working with CommonCrawl data), then you first have to extract the UID (the id column is of type UUID in the databases). Download All files: huggingface-cli download BramVanroy/fineweb-duckdbs --local-dir duckdbs/fineweb/ --include *.duckdb… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/fineweb-duckdbs.1 likes2.1k downloads1y agoHugging Face03samansmink /duckdb_ci_teststextn<1K0 likes929 downloads2y agoHugging Face04ucalyptus /birdbench-duckdb BirdBench Dataset in DuckDB format BirdBench is a benchmark for text-to-SQL capabilities, now available in DuckDB format for improved performance and usability. About BirdBench BirdBench is a comprehensive benchmark dataset for evaluating text-to-SQL capabilities of language models. It features a diverse collection of databases spanning various domains including: Business and finance Entertainment and media Sports and recreation Health and medicine Education Travel and… See the full description on the dataset page: https://huggingface.co/datasets/ucalyptus/birdbench-duckdb.question-answering100M<n<1B2 likes491 downloads2y agoHugging Face05motherduckdb /duckdb-text2sql-25k Dataset Summary The duckdb-text2sql-25k dataset contains 25,000 DuckDB text-2-sql pairs covering diverse aspects of DuckDB's SQL syntax. We synthesized this dataset using Mixtral 8x7B, based on DuckDB's v0.9.2 documentation and Spider schemas that were translated to DuckDB syntax and enriched with nested type columns. Each training sample consists of a natural language prompt, a corresponding (optional) schema, and a resulting query. Each pair furthermore has a category property… See the full description on the dataset page: https://huggingface.co/datasets/motherduckdb/duckdb-text2sql-25k.text10K<n<100K43 likes104 downloads2y agoHugging Face06diicell /duckdb-qa-v3textn<1K0 likes56 downloads8mo agoHugging Face07duckdb-nsql-hub /duckdb-nsql-scorestabularn<1K0 likes54 downloads1y agoHugging Face08huoju /duckdbtest0 likes51 downloads1y agoHugging Face09DavidzzzZZZ /msmarco-duckdb-fts0 likes51 downloads10d agoHugging Face10motherduckdb /duckdb-docbench DocBench: A Synthetic DuckDB Text-to-SQL Benchmark DocBench is a synthetic Text-to-SQL benchmark dataset consisting of 2430 question/sql pairs derived from the DuckDB documentation, specifically designed to probe language models for knowledge of DuckDB-specific SQL functionality. The dataset covers functions, aggregates, operators, statements, keywords, and multi-keyword expressions available in DuckDB 1.1.3 and its default extensions. Dataset Structure Each example… See the full description on the dataset page: https://huggingface.co/datasets/motherduckdb/duckdb-docbench.text1K<n<10K3 likes43 downloads9mo agoHugging Face11diicell /duckdb-dbt-qatextn<1K0 likes43 downloads5mo agoHugging Face12duckdb-nsql-hub /duckdb-nsql-predictionstext1K<n<10K0 likes33 downloads1y agoHugging Face13lhoestq /remote-duckdb0 likes30 downloads3y agoHugging Face14duckdb-nsql-hub /sql-console-prompt SQL Console Text 2 SQL Prompt GitHub Gist Feedback is welcome 🤗. This prompt was based on performance from Qwen on the DuckDB NSQL Benchmark, common dataset types and tasks typical for exploring HF Datasets. This is the prompt used for the Text2SQL inside the SQL Console on Datasets. Example Table Context For the {table_context} we use the SQL DDL CREATE TABLE statement. CREATE TABLE datasets ( "_id" VARCHAR, "id" VARCHAR, "author" VARCHAR… See the full description on the dataset page: https://huggingface.co/datasets/duckdb-nsql-hub/sql-console-prompt.textn<1K9 likes30 downloads2y agoHugging Face15diicell /duckdb-qa-v2textn<1K0 likes26 downloads8mo agoHugging Face16gabrielmbmb /distilabel-duckdb-queries Dataset Card for distilabel-duckdb-queries This dataset has been created with distilabel. The pipeline script was uploaded to easily reproduce the dataset: app.py. It can be run directly using the CLI: distilabel pipeline run --script "https://huggingface.co/datasets/gabrielmbmb/distilabel-duckdb-queries/raw/main/app.py" Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel… See the full description on the dataset page: https://huggingface.co/datasets/gabrielmbmb/distilabel-duckdb-queries.textn<1K0 likes25 downloads2y agoHugging Face17PixelPerv /danbooru-tags-20260518-duckdb Danbooru Tags Analysis (DuckDB Edition) This dataset is an optimized, streamlined derivative of u-haru/danbooru-tags-20260518 (thanks so much for making it available, you rock!). It is specifically engineered to facilitate tag analysis for Illustrious and Anime-focused AI models, described in the Civitai article https://civitai.red/articles/30556. Key Improvements Size Optimization: Retains only a limited subset of columns essential for analytical workloads… See the full description on the dataset page: https://huggingface.co/datasets/PixelPerv/danbooru-tags-20260518-duckdb.10M<n<100M0 likes22 downloads4mo agoHugging Face18rajatm17 /data-jobs-duckdb0 likes22 downloads6d agoHugging Face19duckdb-nsql-hub /duckdb-docstext1K<n<10K4 likes19 downloads2y agoHugging Face20asoria /duckdb_testtextn<1K0 likes17 downloads2y agoHugging Face21ManoharHugs /nyc-taxi-2025-duckdb Taxi-Revenue-and-Surge-Analytics Every taxi ride tells a data story. From raw trip logs to executive insight, analyzed 9.3M NYC Yellow Taxi trips by cleaning messy source data, building tested dimensional models, and uncovering revenue trends, surge patterns, and top-performing zones through an interactive dashboard tabular100K<n<1M0 likes17 downloads6mo agoHugging Face22khaledsayed1 /fineweb-bbc-news-embeddings-DuckDBtext100K<n<1M0 likes16 downloads11mo agoHugging Face23Bhavin1905 /Social-Media-Posts-Dataset-Embeddings-Included-DUCKDB 📊 Social Media Posts Dataset (Embeddings Included) Dataset Description This dataset contains social media posts collected for the purpose of natural language analytics and semantic analysis.It is designed to support trend analysis, topic discovery, sentiment inference, and time-based analytics over historical social media data. The dataset is intended to serve as the data backbone for a natural language analytics system where users can ask questions in plain English and… See the full description on the dataset page: https://huggingface.co/datasets/Bhavin1905/Social-Media-Posts-Dataset-Embeddings-Included-DUCKDB.0 likes16 downloads9mo agoHugging Face24Anshifkk /noaa-gsod-2011-2020-duckdb-bq042 NOAA GSOD 2011-2020 DuckDB baseline Single DuckDB file mirroring the BigQuery public tables bigquery-public-data.noaa_gsod.gsod2011 .. gsod2020 and bigquery-public-data.noaa_gsod.stations one-to-one (same column names, same column types: STRING -> VARCHAR, FLOAT64 -> DOUBLE, INT64 -> BIGINT; no pruning or filtering). Provenance Produced by environment/scripts/rebuild_baseline_from_bigquery.py in the smithy task spider2-lite-bq042-gsod-laguardia-decade, which reads the… See the full description on the dataset page: https://huggingface.co/datasets/Anshifkk/noaa-gsod-2011-2020-duckdb-bq042.0 likes15 downloads6mo agoHugging Face25richie-ghost /dpo_main_finetune_data_duckdbtext10K<n<100K0 likes13 downloads2y agoHugging Face26uckenkare /bb-duckdb-test DuckDB Test Dataset Authorized security testing under HF bug bounty program. textn<1K0 likes13 downloads2mo agoHugging Face27diicell /duckdb-qa-v1textn<1K0 likes9 downloads8mo agoHugging Face28diicell /duckdb-qa-v4textn<1K0 likes9 downloads7mo agoHugging Face29Anshifkk /chembl-29-duckdb-bq430 ChEMBL-29 DuckDB mirror for Spider2-Lite bq430 Offline DuckDB mirror of the four bigquery-public-data.ebi_chembl.* tables referenced by Spider2-Lite task bq430 (matched-molecular-pair deltas): main.activities_29 (18,635,916 rows, 27 VARCHAR cols) main.compound_structures_29 ( 2,084,724 rows, 5 VARCHAR cols) main.compound_properties_29 ( 2,088,293 rows, 23 VARCHAR cols) main.docs_29 ( 81,544 rows, 17 VARCHAR cols) Every column is preserved verbatim… See the full description on the dataset page: https://huggingface.co/datasets/Anshifkk/chembl-29-duckdb-bq430.0 likes9 downloads4mo agoHugging Face30bradex /duckdb0 likes8 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.