datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Survivor
📚 FinePDFs-Edu
350B+ of highly educational tokens from PDFs 📄
What is it?
📚 FinePDFs-Edu dataset consists of 350B+ tokens of educational PDFs filtered from 📄 FinePDFs dataset covering 69 languages.
FinePDFs was created using the formula inspired from FineWeb-Edu, we developed an educational quality classifier using annotations generated by Qwen3-235B-A22B-Instruct-2507 for each of 69 languages present in this dataset.
We then used this classifier to retain only the… See the full description on the dataset page: https://huggingface.co/datasets/Web3Survivor/Survivor.Web3-Dataset
Web3Coders Smart Contracts Dataset – 2026 Edition
Version 2.0 (August 2026) – The largest curated collection of production-grade smart contracts with integrated security audits, gas profiles, vulnerability annotations, and Halal compliance tags for Shariah‑aware Web3 development.
Dataset Description
A comprehensive, multi‑chain dataset of smart contracts (Solidity, Vyper, Rust, and Move) with:
Full source code, bytecode, and ABI
Security audit outcomes (Slither… See the full description on the dataset page: https://huggingface.co/datasets/KurniaKadir/Web3-Dataset.afrofinchain-multilingual-web3
AfroFinChain — Multilingual Web3 & Blockchain Dataset
Multilingual Web3 & blockchain dataset in Yoruba, Hausa, Igbo, and Nigerian Pidgin with 1,451 terminology entries and 1,451 conversational Q&A pairs. Designed for LLM fine-tuning, financial literacy, and conversational AI in low-resource African languages. Uses culturally grounded analogies (e.g., ajo, adashi, isusu) to make DeFi concepts actually understandable.
Built with Adaptive Data by Adaption as part of the Adaption… See the full description on the dataset page: https://huggingface.co/datasets/FirstBML1/afrofinchain-multilingual-web3.Web3-findings-dataset
Smart Contract Audit Findings
This is raw, semi-structured data — not a ready-to-train dataset. It still requires
further cleaning and preparation (deduplication, severity/label normalization, filtering
low-quality or malformed entries, etc.) before it should be used to train or fine-tune an AI model.
A collection of 23,625 smart-contract security audit findings (bug reports), each with a
title, description, proof-of-concept code, recommendation, and severity rating.… See the full description on the dataset page: https://huggingface.co/datasets/0xSojalSec/Web3-findings-dataset.web3-llm-instructions
web3-llm-instructions
Description
web3-llm-instructions is an instruction-following dataset focused on Web3, blockchain, and cryptocurrency concepts.
The dataset is designed for fine-tuning large language models (LLMs) to understand and generate responses about Web3 topics such as DeFi, NFTs, DAOs, smart contracts, and blockchain infrastructure.
Dataset Structure
Each record contains the following fields:
instruction: the task or question
input: optional… See the full description on the dataset page: https://huggingface.co/datasets/YosepMulia/web3-llm-instructions.role-play-bench
Role-play Benchmark
A comprehensive benchmark for evaluating Role-play Agents in Chinese and English scenarios.
Dataset Summary
Role-play Benchmark is designed to evaluate Role-play Agents' ability to deliver immersive role-play experiences through Situated Reenactment. Unlike traditional benchmarks with verifiable answers, Role-play is fundamentally non-verifiable, e.g., there's no single "correct" response when a tsundere character is asked "Do you like me?". Instead… See the full description on the dataset page: https://huggingface.co/datasets/web3w/role-play-bench.indonesian-web3-instruction-50
Indonesian Web3 Instruction Dataset (50 Samples)
This dataset contains 50 Indonesian-language instruction and question–answer pairs focused on Web3, crypto, and blockchain topics.
Use cases
Instruction fine-tuning
Indonesian language modeling
Crypto and Web3 education
Format
JSON Lines with:
instruction
response
