technical
cad-technical-drawings
CAD Technical Drawings, Generated by Cadsy
Turn a STEP model into a labeled technical drawing automatically.
This sample was created with Cadsy from 3D models in the
Zero-to-CAD-100k dataset.
For every STEP model, Cadsy generated:
one drawing using an ASME-style profile;
one drawing using an ISO-style profile; and
structured bounding-box labels for every retained annotation.
That is 65 CAD models, 130 technical drawings and their labels, produced
through one repeatable… See the full description on the dataset page: https://huggingface.co/datasets/cadsy/cad-technical-drawings.reproducibility-datakit-technical-reportDiversity-plot parquet for the gridfm-datakit technical report.
How to reproduce: scripts/datakit_report/configs/README.md on branch genco-paper-repro.
industrial-technical-archive
🚀 Latest Updates (Sep, 2026)
Version: v09.2026 (Verified)
Status: Integrated with 1,000,000+ records.
New Files: product-E-20-09-2026.csv & product-V-20-09-2026.csv.
QTE Technologies: Industrial & Scientific Knowledge Base
Wikidata Entity: Q138411149
IPFS CID: bafybeibogxxuhmzfrsuhcfd4qr4tmc4okhmrcwhp3266hq47ccuyjnjxoq
Official Neural Hub: qtetech.github.io
This is the permanent technical archive for QTE Technologies, ensuring long-term accessibility of… See the full description on the dataset page: https://huggingface.co/datasets/QTE-Technologies/industrial-technical-archive.youtube_technical_v0
youtube_technical_v0
Canonical Phonon ASR dataset. Audio is embedded in Parquet as a Hugging
Face-compatible audio struct with bytes and path.
Hidden eval rows must not be used for training or synthetic prompt generation.
Hub
Dataset id: Infatoshi/youtube_technical_v0 (public)
Companion labels/manifests/attribution: Infatoshi/phonon-youtube-technical
Rows: see summary.json (~118k)
Audio is embedded in Parquet (audio struct: bytes, path)
Split… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/youtube_technical_v0.phonon-youtube-technical
phonon-youtube-technical
This dataset contains Phonon-authored manifests, labels, term/context indexes,
attribution, and local rebuild scripts for technical YouTube speech data.
It intentionally contains no audio files.
Contents
manifests/: segment timing, source URL, video ID, YouTube-reported license,
and audio SHA-256 references.
labels/: Phonon-authored label queues and teacher metadata.
term_context_indexes/: technical term/context indexes used for analysis… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/phonon-youtube-technical.Technical-Architectures-Large
Technical Architectures Large (294k Samples)
Overview
Generating complex, syntactically valid diagram code from natural language requirements is a major challenge for AI models. This dataset bridges that gap by providing over 293,000+ distinct enterprise software architectures generated using two cutting-edge models: GPT-OSS-120B and Qwen3-Coder-Next-FP8.
Unlike simple "toy" examples, these architectures model realistic enterprise systems complete with client… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Technical-Architectures-Large.
