granite
Datasets
All datasets matching “granite”ChartNet
ChartNet: A Million-Scale Multimodal Dataset for Chart Understanding
🌐 Homepage | 📖 arXiv
📝 Changelog
June 3, 2026 — Release of grounded_qa subset and completed reasoning subset (both subject to Notice Regarding Data Availability)
May 15, 2026 — Added link to 30K real-world charts and detailed captions dataset released by our collaborators Abaka AI/2077AI.
April 29, 2026 — Release of an additional 2.5 million row subset core_permissive (subject to… See the full description on the dataset page: https://huggingface.co/datasets/ibm-granite/ChartNet.graniteDr.Sparse-Granite42-8B-eval-b200-otf81-spgemm
Dr.Sparse — Granite 4.2 8B SpGEMM baseline (OTF-81)
ibm-granite/granite-4.2-8b 在 Dr.Sparse OTF 保留测试集上的 SpGEMM baseline。
levels 1-3(排除 level4),单轨迹无树搜索,每矩阵 10 轮迭代,B200 (sm_100)。
结果
level
矩阵
正确
跑赢 cuSPARSE
中位加速比
最大
level1_small
8
1
0
0.244
0.24
level2_medium
33
3
0
0.024
0.62
level3_large
38
4
0
0.092
1.00
合计
79
8
0
0.102
1.00
79 个矩阵里 8 个产出正确 kernel,无一跑赢 cuSPARSE。中位加速比 0.102
表示比 cuSPARSE 慢约十倍;最好的一个仅持平。
编译失败的错误类型分散:cudaMalloc 重载不匹配、const 限定符未去除、… See the full description on the dataset page: https://huggingface.co/datasets/DiogenesChen122/Dr.Sparse-Granite42-8B-eval-b200-otf81-spgemm.GneissWeb
What is it?
Recipe for producing a state-of-the-art LLM pre-training dataset having 10+ Trillion tokens, derived from FineWeb V1.1.0
Evaluation results showing more than 2% avg improvement (with multiple random seeds) over FineWeb V1.1.0 tokens on common benchmarks for a 7B parameter ablation model
Data Prep Kit Notebook for reproducing the annotations and filters on top of FineWeb and Notebook for applying a bloom filter on FineWeb to quickly reproduce an approximate version of… See the full description on the dataset page: https://huggingface.co/datasets/ibm-granite/GneissWeb.wildchat-4.8m_1m_seed1_gemma_granite_metrics_extendedgranite-decisions-synthetic
Granite Decisions synthetic datasets
Original, deterministic English fixtures for Adam Pippert's personal
Granite Decisions project.
The original default config has 162 examples: 54 train, 54 calibration, and 54 test.
These exercise the pipeline; they are not a representative quality benchmark.
Source and license
The source is the project's original template generator, published here as
make_smoke_data.py, from
release v0.1.0,
commit… See the full description on the dataset page: https://huggingface.co/datasets/adampippert/granite-decisions-synthetic.
