datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
profiling-pytorch
Profiling in PyTorch Scripts
Holds all the scripts used for the series "Pofiling in PyTorch"
author_profilinghe corpus for the author profiling analysis contains texts in Russian-language which labeled for 5 tasks:
1) gender -- 13530 texts with the labels, who wrote this: text female or male;
2) age -- 13530 texts with the labels, how old the person who wrote the text. This is a number from 12 to 80. In addition, for the classification task we added 5 age groups: 1-19; 20-29; 30-39; 40-49; 50+;
3) age imitation -- 7574 texts, where crowdsource authors is asked to write three texts:
a) in their natural manner,
b) imitating the style of someone younger,
c) imitating the style of someone older;
4) gender imitation -- 5956 texts, where the crowdsource authors is asked to write texts: in their origin gender and pretending to be the opposite gender;
5) style imitation -- 5956 texts, where crowdsource authors is asked to write a text on behalf of another person of your own gender, with a distortion of the authors usual style.torch-profiling-trace-diffuserspro6-profiling-vaeB.longum_CAZyme_Profiling
Comprehensive Profiling of the Bifidobacterium longum NCC2705 CAZome
Overview
This repository constitutes a robust, automated computational pipeline designed for the deep profiling and structural characterization of the Carbohydrate-Active enZymes (CAZome) repertoire of Bifidobacterium longum strain NCC2705. The analytical framework leverages high-throughput sequence homology and hidden Markov model (HMM) profile alignments against the dbCAN, CAZy, and… See the full description on the dataset page: https://huggingface.co/datasets/QasimHussain/B.longum_CAZyme_Profiling.b200-profiling-vaeexp7-vea-probability-profilingauthor_profilinghe corpus for the author profiling analysis contains texts in Russian-language which labeled for 5 tasks:
1) gender -- 13530 texts with the labels, who wrote this: text female or male;
2) age -- 13530 texts with the labels, how old the person who wrote the text. This is a number from 12 to 80. In addition, for the classification task we added 5 age groups: 1-19; 20-29; 30-39; 40-49; 50+;
3) age imitation -- 7574 texts, where crowdsource authors is asked to write three texts:
a) in their natural manner,
b) imitating the style of someone younger,
c) imitating the style of someone older;
4) gender imitation -- 5956 texts, where the crowdsource authors is asked to write texts: in their origin gender and pretending to be the opposite gender;
5) style imitation -- 5956 texts, where crowdsource authors is asked to write a text on behalf of another person of your own gender, with a distortion of the authors usual style.recipient_profilingtwitter_author_profiling_by_gender_nlpThis dataset was created for a student's Bc work.
The main purpose for which the dataset was created is to use it in author profiling by gender.
Single-Tweet-Per-Author Twitter Dataset
Overview
This dataset consists of Twitter (X) posts with a strict constraint: each author appears exactly once.There is a one-to-one correspondence between tweets and authors.
This design removes author-level accumulation effects and prevents models from exploiting repeated stylistic or… See the full description on the dataset page: https://huggingface.co/datasets/qg2020252627/twitter_author_profiling_by_gender_nlp.moe-routing-profiling-b200so101_audio_profilingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"audio_files_size_in_mb": 100,
"fps": 30,
"splits": {},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/so101_audio_profiling.pharma-combination-state-coherence-profiling-v0.1What this dataset tests
Whether a system can score multi drug combinations
by system level coherence.
The focus is:
cross system alignment
early incoherence flags
net functional shift
Required outputs
combination_coherence_score
dominant_response_basin
cross_system_alignment_vector
early_incoherence_flags
net_functional_shift
Use case
Early screening for combination therapies.
Find coherent combos.
Flag brittle combos before trial spend.
profiling_wikigre
Knowledge Graph-Based Dynamic Factuality Evaluation
Worldview Benchmark for Large Language Models
The authors do not endorse any political, ideological, or moral position implied by the items or by model outputs.
All entries are probes for factual and value-related reasoning and must not be used to profile real people or to justify harmful actions or decisions.
Dataset Description
Profiling: Knowledge Graph-Based Dynamic Factuality Evaluation is a research benchmark… See the full description on the dataset page: https://huggingface.co/datasets/llmpass-ai/profiling_wikigre.silicon-profiling-snapdragon865
Real-Device Silicon Profiling: Snapdragon 865
Per-device inference benchmarks on real Samsung S20 FE 5G phones (Snapdragon 865).
No simulation. Real ARM CPU inference.
Hardware
Property
Value
Chipset
Qualcomm Snapdragon 865 (SM8250)
CPU
Kryo 585: 1x2.84GHz + 3x2.42GHz + 4x1.80GHz
GPU
Adreno 650
NPU
Hexagon Tensor Accelerator
RAM
8GB LPDDR5 (7.47GB total, 3-3.7GB free)
Device
Samsung Galaxy S20 FE 5G (SM-G981V)
Devices connected
39… See the full description on the dataset page: https://huggingface.co/datasets/dispatchAI/silicon-profiling-snapdragon865.TestTimeScalingOlmo-1B-0724-hf-BON_50Q_profilingasia-migration-idps-informal-sites-profiling-and-moveme
Iraq - IDPs Informal sites profiling and movement intentions
Publisher: REACH Initiative · Source: HDX · License: cc-by · Updated: 2025-04-15
Abstract
REACH Initiative supports the humanitarian response in Iraq by conducting assessments of informal sites in Iraq, in partnership with the Camp Coordination and Camp Management (CCCM) cluster. The assessment aims to identify movement intentions and highlight the multi-sectoral needs of internally displaced persons (IDPs)… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-migration-idps-informal-sites-profiling-and-moveme.africa-somalia-somalia-internal-displacement-profiling-in-hargeisa-3c0cc159
Somalia Internal Displacement Profiling in Hargeisa | Africa (Joint IDP Profiling Service (JIPS))
13,265 rows - 1 Africa country/area - 2015-2016 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 13,265 rows from Joint IDP Profiling Service (JIPS), covering Somalia Internal Displacement Profiling in Hargeisa. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-somalia-somalia-internal-displacement-profiling-in-hargeisa-3c0cc159.profiling_baseline2_seqfft_real0_on_real0_seed1000This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 1,
"total_frames": 317,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 1,
"video_files_size_in_mb": 1,
"fps": 15,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/continuallearning/profiling_baseline2_seqfft_real0_on_real0_seed1000.clinical-combination-coherence-profiling-v0.1What this dataset tests
Whether a drug combinationimproves system-wide coherenceor merely suppresses symptoms.
Required outputs
baseline coherence metrics
post-combination coherence metrics
variance and synchrony shifts
net coherence score
interpretation
Use case
Front layer of the Polypharmacy Coherence Matrix.
profilingProfiling-072025-05-06T14-automatic-profilingafrica-sudan-sudan-el-fasher-durable-solutions-profiling-exercise-cc4e948b
Sudan-El Fasher Durable Solutions Profiling Exercise | Africa (Sudan official open data)
22,107 rows - 1 Africa country - 2019 - Repackaged by Electric Sheep Africa
TL;DR
This dataset packages one official XLSX resource from Sudan as
ML-ready Parquet. The source file is the provenance boundary; all usable
indicators or tabular columns from the resource stay together in this repo.
About the source
Source: Sudan-El Fasher Durable Solutions… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-sudan-sudan-el-fasher-durable-solutions-profiling-exercise-cc4e948b.tpu-profilingprofiling-traces
profiling-traces
Perfetto traces hosted as Release assets for sharing via ui.perfetto.dev URL parameter.
reddit_authorship_profiling_romanianclinical-systemic-resilience-gain-profiling-v0.1Resilience-Enhancing Pharmacopeia
Index README
Core premise
Some drugs work across diseases
because they increase system capacity, not because they hit a target.
This collection defines, measures, and deploys that class of drugs.
Not pathology-first.
Resilience-first.
What this pharmacopeia tests
Does a drug broaden the healthy basin
Which fragility axes it buffers
Who should receive it based on systemic vulnerability
These datasets do not ask
“Does this drug treat condition X”
They ask
“Does… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-systemic-resilience-gain-profiling-v0.1.resume-profiling-dataset-llavaprofiling-filtered-data
