datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swiftswiftSwift-OpenX-EmbodimentswiftdataVLM-swiftimage-pointing-1M-sft-swiftswiftdatathe-stack-swift-clean
Dataset 1: TheStack - Swift - Cleaned
Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Swift, a popular statically typed language.
Target Language: Swift
Dataset Size:
Training: 900,000 files
Validation: 50,000 files
Test: 50,000 files
Preprocessing:
Selected Swift as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-swift-clean.swift
Dataset Card for "swift"
More Information needed
Swift-OpenX-Embodiment-wrist-imagesbigquery-swift-unfiltered
GitHub Swift Repositories
Dataset Description
Dataset Summary
This dataset comprises data extracted from GitHub repositories, specifically focusing on Swift code. It was extracted using Google BigQuery and contains detailed information such as the repository name, reference, path, and license.
Source Data
Initial Data Collection and Normalization
The data was collected from GitHub repositories using Google BigQuery. The dataset includes data from… See the full description on the dataset page: https://huggingface.co/datasets/drewparo/bigquery-swift-unfiltered.swift-17
SWIFT 17 — superseded by the n=26 census
SUPERSEDED. This 17-bank press census is superseded by live GET https://councilof.ai/api/swift (n=26 census: 3 LIVE · 9 COMMITTED · 14 DISCOVERED · n_measured=0). Schema notes supersedes: csoai.swift-17/0.1. Prefer csoai/gspc-swift-26 (census mirror) or the live API. Not clients. Not GPI. Not MEASURED.
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
Historical 17 names from Swift… See the full description on the dataset page: https://huggingface.co/datasets/csoai/swift-17.gspc-swift-26
SWIFT census n=26 — mirror of /api/swift
Census mirror of live GET https://councilof.ai/api/swift — n=26 (3 LIVE · 9 COMMITTED · 14 DISCOVERED · n_measured=0). Kind=reader · writes_board=false. Supersedes csoai.swift-17/0.1 / Hub csoai/gspc-swift-17 / csoai/swift-17.
Live SWIFT census: https://councilof.ai/api/swift
Live GSPC board: https://councilof.ai/api/gspc
Honest sourced census of 26 named banks. Not clients. Not GPI. Not MEASURED — no ISO 20022 / copybook / MT artifact… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-swift-26.gspc-swift-50
SWIFT 50 — honesty shell, no public 50
Honest census: live GET https://councilof.ai/api/swift reports n=26 named banks (not 50). Swift cites 40+ in MVP construction; only 26 sourced to a real reachable dated press URL. Remaining ~14+ are not enumerated — no name invented to reach 50. Not clients. Not MEASURED.
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
This slug exists so inbound "swift-50" searches land on an honest… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-swift-50.gspc-swift-17
SWIFT 17 — superseded by the n=26 census
SUPERSEDED by GET https://councilof.ai/api/swift (n=26). Not clients. Not GPI. Not MEASURED (n_measured=0 on the live census). Prefer csoai/gspc-swift-26 or the live API.
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
17 DISCOVERED from Swift press 2026-07-09 was the prior tape. Live census is n=26 (THREE_STATE). This Hub page is a printer/alias, not a second engine. Not a… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-swift-17.silkie-ms-swift
Dataset Card for "silkie-ms-swift"
More Information needed
taylor_swift
Dataset Card for "taylor_swift"
More Information needed
iva-swift-codeint
IVA Swift GitHub Code Dataset
Dataset Description
This is the raw IVA Swift dataset extracted from GitHub.
It contains uncurated Swift files gathered with the purpose to train a code generation model.
The dataset consists of 753693 swift code files from GitHub totaling ~700MB of data.
The dataset was created from the public GitHub dataset on Google BiqQuery.
How to use it
To download the full dataset:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/mvasiliniuc/iva-swift-codeint.mmu_swift_sne_ia
mmu_swift_sne_ia HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_swift_sne_ia.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB can be… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_swift_sne_ia.voxconverse_swift
VoxConverse — Speaker Diarization in the Wild (MS-Swift Format)
This dataset is a reformatted version of VoxConverse for fine-tuning and evaluating multimodal large language models on speaker diarization, packaged in the MS-Swift Parquet format.
Note: The underlying audio is sourced from YouTube videos whose copyright remains with the original owners. This reformatted dataset is intended for research purposes only. For the original annotations and audio, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/AtwMaxime/voxconverse_swift.SWIFTT-bark_beetle_detection_semantic_segmentation
Bark Beetle Detection - Semantic Segmentation Dataset (SWIFTT Project)
📋 Overview
This collection of datasets is devoleped in fullfilment of the research objectives of SWIFTT project (Satellites for Wilderness Inspection and Forest Threat Tracking),
funded by the European Union under Grant Agreement 101082732.
It contains labeled satellite imagery acquired with Sentinel-2 and Sentinel-1, that can be used for developing and evaluating predictive models to detect… See the full description on the dataset page: https://huggingface.co/datasets/AnnalisaAppice/SWIFTT-bark_beetle_detection_semantic_segmentation.PhysicsLENS
License
Our own content (prompts, annotations, manifest) is CC BY-NC-SA 4.0.
Frames retain their source license — see source_license column in
manifest.csv. All sources are included as files except EgoDex
(CC BY-NC-ND 4.0, which prohibits redistributing derivatives); those
2 frames are marked included: no in the manifest.
Source
License
Humanoid Everyday
Apache-2.0
Unitree G1 Dex1
Apache-2.0
Unitree G1 Dex3
Apache-2.0
Unitree Z1 dual-arm
Apache-2.0
RoboCOIN… See the full description on the dataset page: https://huggingface.co/datasets/swiftrando/PhysicsLENS.swift_testSwift-OpenX-Embodiment-action-chunk-jsonsswiftplan-isaac-sim
SwiftPlan Isaac Sim Dataset
This dataset contains Isaac Sim observation images for frame-level high-level action selection in robotic task planning.
Each sample includes:
an RGB observation image,
a task instruction,
a frame-level high-level action label,
an action type,
an optional target object.
The dataset is designed for execution-time high-level decision making, where a model selects the next high-level action from the current observation and task instruction.… See the full description on the dataset page: https://huggingface.co/datasets/Kuoskyler/swiftplan-isaac-sim.LLaVA-Video-small-swift
Dataset Card LLaVA-Video-small-swift
Small subset of LLaVA-Video-178K for educational purposes to learn how to fine-tune video models.
iva-swift-codeint-clean-train
IVA Swift GitHub Code Dataset
Dataset Description
This is the curated train split of IVA Swift dataset extracted from GitHub.
It contains curated Swift files gathered with the purpose to train a code generation model.
The dataset consists of 320000 Swift code files from GitHub.
Here is the unsliced curated dataset and
here is the raw dataset.
How to use it
To download the full dataset:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/mvasiliniuc/iva-swift-codeint-clean-train.swift-reasoning-rollouts-deepscaler-ministral8b
DeepScaleR Reasoning Rollouts (Ministral-8B)
This dataset contains reasoning rollouts used to train the SWIFT reward head.
Paper page: https://huggingface.co/papers/2505.12225
GitHub: https://github.com/aster2024/SWIFT/
Generator model: mistralai/Ministral-8B-Instruct-2410 (https://huggingface.co/mistralai/Ministral-8B-Instruct-2410)
Dataset Description
This dataset contains 10000 samples corresponding to the Generalization Test setup.
Source: DeepScaleR.
Generator:… See the full description on the dataset page: https://huggingface.co/datasets/Aster2024/swift-reasoning-rollouts-deepscaler-ministral8b.bigquery-swift-filtered-no-duplicate
Dataset Card for "bigquery-swift-unfiltered-no-duplicate"
More Information needed
rlvr-code-data-swift-edited
