datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SemanticSeg
Dataset Card for SemanticSeg
This semantic segmentation dataset introduced in the paper Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation.
This dataset is used to train the segmenter.
Dataset Details
Dataset Description
SemanticSeg contains around 16 segmentation categories, with each category containing at least 2k instances. The varying cut rates across categories can also help the segmenter learn… See the full description on the dataset page: https://huggingface.co/datasets/Syon-Li/SemanticSeg.Semantic-Search-Engine-with-Vectorized-DB
Semantic Search Engine with Vectorized DB — Artifacts
This repository hosts the pre-computed on-disk index artifacts for the 20,000,000 vector database (OpenSubtitles_en_20M_emb_64.dat), built for the Advanced Database Systems project (Cairo University, Faculty of Engineering).
📁 Repository Structure
semantic-search-artifacts/
│
├── README.md # Repository documentation & usage guide
│
├── production/
│ ├── m1_ivf_k4096/… See the full description on the dataset page: https://huggingface.co/datasets/Final-Progs/Semantic-Search-Engine-with-Vectorized-DB.semantic-scholar-scraper
Semantic Scholar Scraper · Papers, Authors, Citations & Venues
Scrape academic research papers, authors, citations, venues, and open-access metadata from Semantic Scholar API. Features rate-limit backoff resilience and pay-per-event pricing.
Rows in this dataset
450
Fields
19
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/semantic-scholar-scraper.
