datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
luatdo-graph
luatdo-graph
A knowledge graph over Vietnamese law, built from about 128,000 documents by luatdo.
This is the result of running the pipeline, published so that nobody has to run it again.
The pipeline takes days and several hundred dollars of model calls, and the output is the same for everyone.
What is in it
Nodes
8,175,346
Relationships
9,119,011
Node tables
14
Relationship tables
19
Node labels
13
Parquet
623MB across 47 files
Neo4j… See the full description on the dataset page: https://huggingface.co/datasets/open-index/luatdo-graph.Stack2Graph_VD_lua
LUA StackOverflow Vector Dataset
Summary
This Hugging Face dataset repository contains the LUA shard of the Stack2Graph vector-database component as restorable Qdrant artifacts plus portable Parquet fallback files.
Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper.
The artifacts are intended for semantic and hybrid retrieval, graph entry-point finding, and… See the full description on the dataset page: https://huggingface.co/datasets/Mo7art/Stack2Graph_VD_lua.
