Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01abisee /cnn_dailymail Dataset Card for CNN Dailymail Dataset Dataset Summary The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering. Supported Tasks and Leaderboards 'summarization': Versions… See the full description on the dataset page: https://huggingface.co/datasets/abisee/cnn_dailymail.textsummarization100K<n<1M353 likes66k downloads3y agoHugging Face02abigailasanchez4139 /helios0 likes16k downloads2h agoHugging Face03abigailhaddad /legislative-issue-tracker Legislative Issue Tracker Bills, legislative actions, floor speeches, hearings, and committee reports that touch a specific federal statute — currently the Paperwork Reduction Act (44 U.S.C. ch. 35, subch. I) — with every mention classified as amends, exempts, references, or related. The distinction is the point: Congress amends the PRA rarely (254 bills) but exempts individual programs from it constantly (845 bills). Built by abigail-64/legislative-issue-tracker. Everything… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/legislative-issue-tracker.textn<1K0 likes6.8k downloads7h agoHugging Face04abigailhaddad /foia-reading-room-documents Foia Reading Room Documents Documents from federal FOIA reading rooms and Inspector General report libraries: audits, inspections, investigative summaries and records released under the Freedom of Information Act. Every document here was published by a US federal agency and is a work of the United States government. Nothing has been altered: files are byte-identical to what the agency posted, and the checksum in metadata.parquet is of the original bytes. Why this… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/foia-reading-room-documents.text-retrieval0 likes6.5k downloads28d agoHugging Face05abigailhaddad /usaspending-bulk-awards USAspending bulk awards — contracts & assistance Clean, partitioned, query-ready Parquet mirror of the public USAspending Award Data Archive (prime contract and financial-assistance transactions, FY2007–present, all agencies). The source publishes 4,600 per-agency ZIP/CSV files (830 GB uncompressed). This dataset normalizes them to typed, zstd-compressed Parquet (~8× smaller) with amount columns as double and date columns as date, partitioned for fast predicate-pushdown… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/usaspending-bulk-awards.100M<n<1B0 likes3.1k downloads5d agoHugging Face06Abirate /english_quotes Dataset Card for English quotes I-Dataset Summary english_quotes is a dataset of all the quotes retrieved from goodreads quotes. This dataset can be used for multi-label text classification and text generation. The content of each quote is in English and concerns the domain of datasets for NLP and beyond. II-Supported Tasks and Leaderboards Multi-label text classification : The dataset can be used to train a model for text-classification, which consists of… See the full description on the dataset page: https://huggingface.co/datasets/Abirate/english_quotes.texttext-classification1K<n<10K109 likes2.7k downloads4y agoHugging Face07abigailhaddad /usajobs-scraping USAJOBS announcement text The full text of federal job announcements, scraped from usajobs.gov and joined to the structured fields from the USAJOBS Historical API. About 3.2 million announcements from September 2013 through September 2026, updated daily. Why this exists The USAJOBS API is a poor source for announcement text, in two ways. The Search API only lists jobs that are open right now, so anything that opens and closes between two collection runs is never… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/usajobs-scraping.tabulartext-classification1M<n<10M0 likes2.6k downloads9h agoHugging Face08adamliewehr /512x1_ABI_CloudSat0 likes2k downloads2mo agoHugging Face09raincandy-u /Ab-Integro Ab Integro Ab Integro 是面向《欧陆风云 IV》1.37.5 的综合模组。默认内容为简体中文;英文以独立覆盖层提供。 内容 原版与整合内容的简体中文本地化 界面、旗帜、地图模式、地形和单位标记调整 外交、事件、警报和操作体验扩展 Comprehensive Map 与 India Extended 的地图和历史内容 加载界面名人名言提示(67 条考据后收录) 加载界面油画(1400–1800 年公版作品,来源见 ATTRIBUTION.md) 完整的版本记录与来源清单见 CHANGELOG.md。 安装 将 Ab_Integro 模组目录和对应的 .mod 描述文件放入 EU4 用户模组目录,在启动器中启用 Ab Integro。 中文显示需要 EU4 双字节汉化补丁。 需要英文时,先启用 Ab Integro,再启用 Ab Integro English。英文覆盖层只回填语言文本,不能单独启用。 致谢… See the full description on the dataset page: https://huggingface.co/datasets/raincandy-u/Ab-Integro.0 likes1.7k downloads2mo agoHugging Face10abigailramirez79652 /rush0 likes1.4k downloads2h agoHugging Face11abigailhaddad /dod-daily-contracts DoD daily contract announcements Scraped from war.gov/News/Contracts, one row per contract award (not per day). column date announcement date (ISO 8601), from the article URL agency branch/agency header the award was listed under (ARMY, NAVY, AIR FORCE, DEFENSE LOGISTICS AGENCY, ...), as published -- typos and casing are the source's own, not normalized company best-effort extraction of the awardee's name from the start of text (~95% match rate; None when not… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/dod-daily-contracts.text10K<n<100K1 likes1k downloads22h agoHugging Face12abigailturner53352 /chameleon0 likes953 downloads2h agoHugging Face13abideen /pretrain_corpustext10M<n<100M2 likes759 downloads3y agoHugging Face14abigailhaddad /sam-solicitation-documents Sam Solicitation Documents Attachments from federal solicitation notices on SAM.gov: statements of work, performance work statements, justifications, amendments, wage determinations and the rest of the paperwork that accompanies a federal contract opportunity. Every document here was published by a US federal agency and is a work of the United States government. Nothing has been altered: files are byte-identical to what the agency posted, and the checksum in metadata.parquet is… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/sam-solicitation-documents.text-retrieval0 likes736 downloads23d agoHugging Face15Abin0008 /real-infrared-maritime-vessel-dataset Real Infrared Maritime Vessel Dataset Real infrared imagery of maritime vessels. The dataset is provided in three forms — full-frame detection images, per-object classification crops, and a hand-curated subset. Classes (7): liner, bulk carrier, warship, sailboat, canoe, container ship, fishing boat. Layout real-infrared-maritime-vessel-dataset/ ├── original/ Full-frame IR images + XML bounding-box labels (detection) │ ├── images/{train,test}/*.jpg… See the full description on the dataset page: https://huggingface.co/datasets/Abin0008/real-infrared-maritime-vessel-dataset.imageimage-classification10K<n<100K0 likes601 downloads3mo agoHugging Face16abigailhaddad /federal-public-lands-spending Federal Public Lands Spending No longer updated. The automated sync stopped on 2026-09-25. The data here reflects the USAspending Award Data Archive as of the last update on 2026-09-14. If you need it kept current, get in touch. Contract and grant transaction data from the USAspending Award Data Archive for federal agencies that manage public lands. Interactive demo: Pipeline code: github.com/abigailhaddad/federal-public-lands-contracting Agencies covered… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/federal-public-lands-spending.tabulartabular-classification1M<n<10M1 likes571 downloads15d agoHugging Face17abir-hr196 /clt_gpt2_tokenized_control Fresh multilingual GPT-2 CLT control data Sequential, unshuffled control sample for CLT null experiments. For each language, complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision 0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded (including the complete document that crossed the threshold), after which complete documents were retained until at least 100,000,000 tokens were collected. Data are… See the full description on the dataset page: https://huggingface.co/datasets/abir-hr196/clt_gpt2_tokenized_control.texttext-generation100K<n<1M0 likes543 downloads3mo agoHugging Face18jason1966 /abidhussai512_historical-ai-chip-stock-prices-20182026 Historical AI Chip Stock Prices 2018–2026 Historical AI Chip Stock Prices (AMD, ASML, INTC, NVDA) 2018–2026 Dataset Info Source: Kaggle Original Size: 0.26 MB Kaggle Downloads: 14 Files: 1 Files ai_chip_stocks_2018_2026.csv Mirrored from Kaggle text1K<n<10K0 likes472 downloads6mo agoHugging Face19abigailhaddad /hhs-dab-decisions HHS Departmental Appeals Board decisions Every decision the HHS Departmental Appeals Board has published, as typed Parquet with the header and parts of the body parsed into columns. The Board is an administrative tribunal inside HHS. Its ALJs (Civil Remedies Division) hear a case first and its Appellate Division reviews them, so the two files are two levels of the same tribunal and can be joined on reviews_decision_no / appealed_in. file tribunal decisions span… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/hhs-dab-decisions.tabular10K<n<100K0 likes313 downloads1d agoHugging Face20Abirami /tamilwikipediadatasetannotations_creators: found language: Tamil language_creators: found license: [] multilinguality: multilingual pretty_name: tamilwikipediadataset size_categories: 100K<n<1M source_datasets: [] tags: [] task_categories: summarization task_ids: [] text100K<n<1M2 likes311 downloads4y agoHugging Face21nyu-dice-lab /lm-eval-results-abideen-AlphaMonarch-daser-private Dataset Card for Evaluation run of abideen/AlphaMonarch-daser Dataset automatically created during the evaluation run of model abideen/AlphaMonarch-daser The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-abideen-AlphaMonarch-daser-private.tabular100K<n<1M0 likes256 downloads2y agoHugging Face22Abirate /french_book_reviews Dataset Card for French book reviews I-Dataset Summary The majority of review datasets are in English. There are datasets in other languages, but not many. Through this work, I would like to enrich the datasets in the French language(my mother tongue with Arabic).The data was retrieved from two French websites: Babelio and Critiques LibresLike Wikipedia, these two French sites are made possible by the contributions of volunteers who use the Internet to share their… See the full description on the dataset page: https://huggingface.co/datasets/Abirate/french_book_reviews.tabulartext-classification1K<n<10K8 likes244 downloads4y agoHugging Face23AbirAshraf51611 /waltoncolorimagen<1K0 likes226 downloads2mo agoHugging Face24jason1966 /abidhussai512_us-tech-and-ev-stock-market-dataset-2018-2026 📈 US Tech & EV Stock Market Dataset (2018 – 2026) US Tech & EV Stock Market Dataset of AAPL, MSFT & TSLA using Python API Dataset Info Source: Kaggle Original Size: 0.20 MB Kaggle Downloads: 381 Files: 1 Files stock_market_data.csv Mirrored from Kaggle text1K<n<10K1 likes216 downloads6mo agoHugging Face25jessenorthup24 /abiblicallens-study-audio Tyndale Open Study Notes, read aloud The Tyndale Open Study Notes material used in the free Bible app A Biblical Lens, read aloud by an AI voice. In TYNDALE/am_michael/: notes/<book>/<chapter>.mp3: the study notes that begin in each chapter. notes/<book>.json gives each note's start and end, {"<chapter>": [[note index, start, end], …]} in seconds, where the index is the note's position in the app's studynotes/<book>.json. articles/<book>/<i>.mp3: each theme article.… See the full description on the dataset page: https://huggingface.co/datasets/jessenorthup24/abiblicallens-study-audio.audiotext-to-speech1K<n<10K0 likes204 downloads2d agoHugging Face26jessenorthup24 /abiblicallens-audio A Biblical Lens audio Bibles (Kokoro-82M) Whole public-domain Bibles, read aloud chapter by chapter for the free Bible app A Biblical Lens. Folder Bible Voice WEB/af_heart/ World English Bible (WEB) Heart (af_heart, American) BSB/bm_george/ Berean Standard Bible (BSB) George (bm_george, British) In each folder: <book>/<chapter>.mp3: each chapter, announced ("Genesis, chapter 1."), mono MP3, 24 kHz, 48 kbps. <book>.json: when each verse starts and ends in its… See the full description on the dataset page: https://huggingface.co/datasets/jessenorthup24/abiblicallens-audio.audiotext-to-speech1K<n<10K0 likes201 downloads1d agoHugging Face27AbiralArch /hardware-cvdp-complete CVDP - Comprehensive Verilog Design Problems (Complete Dataset) 🎯 782 out of 783 problems from the official CVDP benchmark by NVIDIA Research 🔥 Dataset Overview This is the most complete version of the Comprehensive Verilog Design Problems (CVDP) benchmark available, containing 782 problems across 13 task categories. CVDP is designed to evaluate Large Language Models and agents on RTL design and verification tasks. 📊 Dataset Statistics Total Problems: 772… See the full description on the dataset page: https://huggingface.co/datasets/AbiralArch/hardware-cvdp-complete.text-generation1K<n<10K1 likes193 downloads1y agoHugging Face28abinzzz /ForeLen Dataset Summary ForeLen is a comprehensive benchmark designed to evaluate Large Language Model (LLM) output length prediction. It includes long-sequence, Chain-of-Thought (CoT), and reinforcement learning (RL) sampling data, enabling the community to rigorously test both static and dynamic length predictors. 🗂 Data Structure Data is organized by model and scenario: Model Scenarios Splits Llama3.2 1B, 3B LongSeq, Reasoning, RL train, validation, test Qwen2.5… See the full description on the dataset page: https://huggingface.co/datasets/abinzzz/ForeLen.text100K<n<1M3 likes187 downloads8mo agoHugging Face29AbijahKaj /kicad-netlist-sft-dataset KiCad Netlist SFT Dataset Training dataset for fine-tuning LLMs to generate valid KiCad electronic circuit netlists from natural language descriptions. Contains 100,179 examples with two complementary output formats: Blog post: Teaching a Small LLM to Design Electronic Circuits: Fine-Tuning Qwen3-4B on 100K KiCad Netlists Format Examples Description SKiDL Python 100,179 Executable Python netlists in the messages assistant field Structured JSON 100,179 Parallel… See the full description on the dataset page: https://huggingface.co/datasets/AbijahKaj/kicad-netlist-sft-dataset.texttext-generation100K<n<1M1 likes179 downloads22d agoHugging Face30abidlabs /test-translation-datasettextn<1K0 likes178 downloads5y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.