datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arb-officeqa-data-packet
ARB OfficeQA data packet
The training data packet of Applied RSI Bench's (ARB) OfficeQA benchmark: the documents a
post-training agent may study before it trains a model for OfficeQA, Databricks' questions on
the U.S. Treasury Bulletin (1939-2025). One archive, L.tar.gz, built by ARB's data_packet
pipeline (map, scout, script, list, build) from OfficeQA's questions without their answers,
kept to sources whose terms allow redistribution, plus OfficeQA's own corpus.… See the full description on the dataset page: https://huggingface.co/datasets/junlinw/arb-officeqa-data-packet.realigns-business-open-data-pack
Realigns Business Open Data Pack
A developer-friendly open business dataset pack prepared by Realigns Inc. for AI assistants, dashboards, market research tools, RAG systems, financial research tools, HR analytics, and business intelligence applications.
This repository combines multiple public business and economic datasets into clean CSV and JSONL formats.
Dataset Index
For a complete folder-by-folder overview of all included datasets, see:
DATASETS.md… See the full description on the dataset page: https://huggingface.co/datasets/realigns/realigns-business-open-data-pack.Packora-data
Packora data
Project links
Project page
Try Packora (live demo)
Paper
Code
Model checkpoints
Hugging Face collection
This repository contains the CSV refcode manifests used to reconstruct the
Packora training and benchmark datasets with a licensed Cambridge Structural
Database installation. It also provides the aggregate dataset statistics used
by the released models. It does not contain CSD structures or prebuilt LMDBs.
csv_manifests/csd.csv is the Packora… See the full description on the dataset page: https://huggingface.co/datasets/nayoung10/Packora-data.arb-fav2-data-packet
fav2 data packet
Study material for the fav2 benchmark of Applied RSI Bench (ARB): the documents an analyst would read to answer the 27 public Vals AI Finance Agent v2 questions (SEC filings of the companies the questions name and of the wider US market, federal agency pages and data, regulations), as
plain text where possible. An ARB post-training agent gets the packet read-only at
/home/agent/workspace/data_packet/ before its model is evaluated.
This is a redistributable… See the full description on the dataset page: https://huggingface.co/datasets/junlinw/arb-fav2-data-packet.skirmish-datapacks
Skirmish data packs
Private data packs consumed by the Skirmish
continuous-integration system tests. They are produced by the
gameplay_export tool in cornerian/skirmish-assets from a verified
Super Smash Bros. Melee (USA 1.02) disc extraction, using the layouts of the
pinned doldecomp/melee revision named in each pack's manifest.json.
Path
Contents
gameplay/v1/skirmish-gameplay-v1.tar.gz
Minimal headless gameplay data: Fox (attributes, bones, hurtboxes, ECB, jab) and… See the full description on the dataset page: https://huggingface.co/datasets/cornerian/skirmish-datapacks.arb-harvey-lab-data-packet
Harvey LAB data packet
Study material for the harvey-lab benchmark of Applied RSI Bench (ARB): the
documents a lawyer would read to do the work in
Harvey LAB's tasks (statutes,
regulations, agency guidance, court opinions, filings, model documents and
practice material), as plain text where possible. An ARB post-training agent gets
the packet read-only at /home/agent/workspace/data_packet/ before its model is
evaluated.
The packet holds no LAB scenarios, case files, task… See the full description on the dataset page: https://huggingface.co/datasets/junlinw/arb-harvey-lab-data-packet.packaged-snack-nutrition-data
24-679 (Fall 2026): Packaged Snack Nutrition
shanexf/packaged-snack-nutrition-data
Nutrition Facts values for 30 packaged snack products, hand-collected from package labels, plus
explicitly marked synthetic variants. The classroom task is multiclass classification: predict the
snack category (chips, crackers, cookies, candy, granola_bars) from eight per-serving nutrition numbers.
Purpose
Built for the 24-679 (Fall 2026, Carnegie Mellon University) assignment on… See the full description on the dataset page: https://huggingface.co/datasets/shanexf/packaged-snack-nutrition-data.realigns-business-open-data-pack
Realigns Business Open Data Pack
A developer-friendly open business dataset pack prepared by Realigns Inc. for AI assistants, dashboards, market research tools, RAG systems, financial research tools, HR analytics, and business intelligence applications.
This repository combines multiple public business and economic datasets into clean CSV and JSONL formats.
Dataset Index
For a complete folder-by-folder overview of all included datasets, see:
DATASETS.md… See the full description on the dataset page: https://huggingface.co/datasets/jamesxu1970/realigns-business-open-data-pack.certified-packings-data
Certified circle packings (sum of radii) and autocorrelation step functions
Author: Elian Alfonso López Preciado, Independent Researcher, León, Guanajuato, México. Code: https://github.com/elianalfonsolopezpreciado/certified-packings (MIT). This dataset (certificates, .pck files, tables): CC-BY-4.0.
What is in here
certificates/ - JSON certificates. A_n<k>.json / sum_radii_n<k>.json: packings of n disjoint circles in the unit square [0,1]^2 (circles: [[x, y, r]… See the full description on the dataset page: https://huggingface.co/datasets/elianalfonsolopezpreciado/certified-packings-data.tourism-package-prediction-dataCLIP-Cross-Attn-MUX-Training-Data-Pack
CLIP-MUX Training Data Pack
This repository is the data/metadata companion for reproducing the CLIP-MUX / x-attention CLIP training pipeline.Used to train model: zer0int/CLIP-ViT-L-14-Cross-Attn-Read-NoRead-ModeMUX
This is not one monolithic dataset under one license. Each component is independently scoped and carries its own LICENSE or NOTICE file.
The pack intentionally separates:
assets that can be redistributed directly;
runtime labels/manifests derived from upstream… See the full description on the dataset page: https://huggingface.co/datasets/zer0int/CLIP-Cross-Attn-MUX-Training-Data-Pack.arb-bigfinancebench-data-packet
BigFinanceBench data packet
Study material for the bigfinancebench benchmark of Applied RSI Bench (ARB): the documents an analyst would read to answer the 50 public BigFinanceBench questions (SEC filings and structured data, SEC rules, IRS and Treasury pages, federal statistics), as
plain text where possible. An ARB post-training agent gets the packet read-only at
/home/agent/workspace/data_packet/ before its model is evaluated.
This is a redistributable build. It keeps only… See the full description on the dataset page: https://huggingface.co/datasets/junlinw/arb-bigfinancebench-data-packet.arb-diligencebench-data-packet
DiligenceBench data packet
Study material for the diligencebench benchmark of Applied RSI Bench (ARB): the documents an analyst would read to write the 150 DiligenceBench due-diligence memos (SEC filings, bank regulators' guidance and data, federal agency pages and datasets), as
plain text where possible. An ARB post-training agent gets the packet read-only at
/home/agent/workspace/data_packet/ before its model is evaluated.
This is a redistributable build. It keeps only… See the full description on the dataset page: https://huggingface.co/datasets/junlinw/arb-diligencebench-data-packet.tourism-package-datatourism-package-prediction-datatourism-package-datatourism-package-datatest-data-pack-0a5e02dc
Test data pack 0a5e02dc
Created by tool orchestrator test UI
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by the sync operator.… See the full description on the dataset page: https://huggingface.co/datasets/physicl-test/test-data-pack-0a5e02dc.data_packproduct_packaging_brand_training_data_v1.0vbpl-datapack
VBPL — văn bản pháp luật Việt Nam
101.356 văn bản quy phạm pháp luật Việt Nam, dữ liệu tới 2026-08-03,
mỗi văn bản một dòng (đã dedupe theo số ký hiệu).
Nguồn: Crawl từ vbpl.vn — CSDL quốc gia về văn bản
pháp luật của Bộ Tư pháp.
Cột
id · so_ky_hieu · tieu_de · loai_van_ban · co_quan_ban_hanh ·
ngay_ban_hanh · ngay_hieu_luc · tinh_trang · pham_vi · nganh ·
linh_vuc · ngay_het_hieu_luc · url · noi_dung
Ngày ở dạng ISO YYYY-MM-DD. noi_dung là toàn văn… See the full description on the dataset page: https://huggingface.co/datasets/billpham/vbpl-datapack.12-Handbags-Jewelry-Sample-Pack50-Enterprise-CLEAR-Compliant-POC
👜 BWS Luxury Goods: Compliance-Native Multimodal Tokens (POC)
🛡️ Engineering Evaluation Sandbox (Active 7-Day Access)
Technical Ingestion Portal: s3://createphotos (Whitelisted buckets only)
Secure Evaluation Link: Download 12_Handbags_Jewelry_Sample_Pack_50_Enterprise_POC.zip
Direct Manifest Auditor: BWS Forensic Manifest Repository
Procurement: All assets are 2026 US CLEAR Act compliant. Access is granted to whitelisted engineering nodes only. Forward your AWS Account ID… See the full description on the dataset page: https://huggingface.co/datasets/BWS-Data-Solutions/12-Handbags-Jewelry-Sample-Pack50-Enterprise-CLEAR-Compliant-POC.opencode-public-data-pack-docker_input2-22-renders-512-20260612t153851z
OpenCode Public Data Pack docker_input2 22 renders 512 20260612T153851Z
Public data pack created from docker_input2.json with 22 renders at 512x512.
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are… See the full description on the dataset page: https://huggingface.co/datasets/physicl-test/opencode-public-data-pack-docker_input2-22-renders-512-20260612t153851z.tourism-package-datatourism-package-prediction-dataRobot_humanml_data_packedtourism-package-prediction-dataopencode-public-data-pack-docker_input1-20-renders-512-20260612t125042z
OpenCode Public Data Pack docker_input1 20 renders 512 20260612T125042Z
Public data pack created from docker_input1.json with 20 renders at 512x512.
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are… See the full description on the dataset page: https://huggingface.co/datasets/physicl-test/opencode-public-data-pack-docker_input1-20-renders-512-20260612t125042z.test-data-pack-dc38769e-68996bd2
Test data pack dc38769e
Created by tool orchestrator test UI
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by the sync operator.… See the full description on the dataset page: https://huggingface.co/datasets/physicl-test/test-data-pack-dc38769e-68996bd2.tourism-package-data
