datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MTBLS289
MTBLS289
A dataset of ~110 paired Whole Slide Images (WSI) and Mass Spectrometry Images (MSI).
Publication: Gerbig, S., Golf, O., Balog, J. et al. Analysis of colorectal adenocarcinoma tissue by desorption electrospray ionization mass spectrometric imaging. Anal Bioanal Chem 403, 2315–2325 (2012).
indonesia-macro-ohlcv-1990-2026
📊 Dataset: AryaP04/indonesia-macro-ohlcv-1990-2026
Indonesia Cross-Asset & Macroeconomic Historical Dataset (1990–2026)Format: Comma-Separated Values (CSV), UTF-8Periode Waktu: 2 Januari 1990 s.d. 27 Mei 2026 (9.486 Baris Data Harian)
📌 1. Ikhtisar Dataset (Overview)
Dataset AryaP04/indonesia-macro-ohlcv-1990-2026 (merged_macro_ohlcv.csv) adalah dataset deret waktu harian (daily time-series) yang menggabungkan pergerakan harga komoditas utama dunia, nilai… See the full description on the dataset page: https://huggingface.co/datasets/AryaP04/indonesia-macro-ohlcv-1990-2026.aryashreeshrestha_single-labeled-news-dataset-for-classification
Single Labeled News Dataset for Classification
Mirror of the Kaggle dataset aryashreeshrestha/single-labeled-news-dataset-for-classification by Arya Shree Shrestha, released under MIT. All credit goes to the original author; please cite and link the Kaggle page when using this data.
License
MIT License, Copyright (c) Arya Shree Shrestha. The full license text is in LICENSE and applies to all files in this repository.
Original description (from… See the full description on the dataset page: https://huggingface.co/datasets/haoxianc/aryashreeshrestha_single-labeled-news-dataset-for-classification.processed_dataset_orca-math-word-problems-200kDataset Description:
This dataset contains data that has undergone two preprocessing steps:
Removal of Instructions with Less Than 100 Tokens in Response: Instructions with less than 100 tokens in the response have been removed from the dataset. This preprocessing step helps to ensure that the dataset contains substantial and informative responses.
Data Deduplication by Grouping Using Cosine Similarity (Threshold > 0.95): Data deduplication has been performed by grouping similar instances… See the full description on the dataset page: https://huggingface.co/datasets/AryanAnuj/processed_dataset_orca-math-word-problems-200k.english-marathiHindu-Gods-DB
Hindu Gods Names Dataset
This dataset contains a curated list of names and titles attributed to various Hindu deities, along with their meanings.
It is intended for use in natural language processing, devotional content generation, cultural studies, and more.
Filename: Hindu-Gods-DB.csv
Format: CSV
Encoding: UTF-8
Columns
Column Name
Description
Deity
The Hindu god or goddess associated with the name.
Name
A name or title used for that deity.
Meaning
The… See the full description on the dataset page: https://huggingface.co/datasets/aryansai/Hindu-Gods-DB.piimask-hackathon
mysqlclient
This project is a fork of MySQLdb1.
This project adds Python 3 support and fixed many bugs.
PyPI: https://pypi.org/project/mysqlclient/
GitHub: https://github.com/PyMySQL/mysqlclient
Support
Do Not use Github Issue Tracker to ask help. OSS Maintainer is not free tech support
When your question looks relating to Python rather than MySQL:
Python mailing list python-list
Slack pythondev.slack.com
Or when you have question about MySQL:
MySQL Community on… See the full description on the dataset page: https://huggingface.co/datasets/arychaud/piimask-hackathon.headlines-bias
Political News Headlines Bias Dataset
This dataset contains 1,000 synthetic political news headlines labeled with perceived political bias:
Left
Center
Right
Columns:
headline: The news headline
bias: The associated political bias label
This dataset was created for the BiasLens project to train LLMs and classifiers for detecting political bias in media.
irumozhiIruMozhi is a human-translated dataset of parallel text in Literary and
Spoken Tamil, using sentences taken from Wikipedia. For more details, see the
paper.
english-to-teluguThis dataset contains translations of English sentences to multiple Indian languages.
recipesAryl-Halides-Source-DataThis is the full data set for the aryl halide benchmark.
yumcraftdescription_to_claims_split709_parser_datasetparser_dataset_sharegptcs661-big-data-project-datasettourism-wellness-package-datasetopenness_instructionenglish-to-tamilshipdataset709_reasoned_parser_datasetcredit_card_fraud-detectiontest_dataenglish-to-marathi855_parser_dataset855_reasoned_parser_dataset897_reasoned_parser_datasetaryaumeshlphilly-crime-2006-2025
