datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Resume-Analysis-CoTR
Resume Reasoning and Feedback Dataset
Dataset Description
This dataset contains approximately 417 examples designed to facilitate research and development in automated resume analysis and feedback generation. Each data point consists of a user query regarding their resume, a simulated internal analysis (chain-of-thought) performed by an expert persona, and a final, user-facing feedback response derived solely from that analysis.
The dataset captures a two-step reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Akhil-Theerthala/Resume-Analysis-CoTR.telco-analysis
telco-analysis
Aggregate (not per-question) evaluation results for my MSc thesis (GSMA Open Telco AI Initiative)
— RAG configuration comparison across 6 generator models, built for supervisor review. Source data
lives in LiamDuero/telco-eval; this repo
is a curated, flattened extract of just the final numbers (accuracy, cost, latency, retrieval
metrics, etc.) with the large per-question arrays stripped out.
Built by telcolens/liam/build_telco_analysis_repo.py, re-runnable at any… See the full description on the dataset page: https://huggingface.co/datasets/LiamDuero/telco-analysis.VisOnlyQA_eval_analysis_6
VisOnlyQA
🌐 Project Website | 📄 Paper | 🤗 Dataset | 🔥 VLMEvalKit
This repository contains the code and data for the paper "VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information".
VisOnlyQA is designed to evaluate the visual perception capability of large vision language models (LVLMs) on geometric information of scientific figures. The evaluation set includes 1,200 mlutiple choice questions in 12 visual perception tasks on 4… See the full description on the dataset page: https://huggingface.co/datasets/ryokamoi/VisOnlyQA_eval_analysis_6.adaption-si-dimensional-analysis-12k
SI Dimensional Analysis
Dimensional reasoning in SI base units: is a relation dimensionally consistent, reduce a combination to base units, find the term that cannot belong in a sum, and decide whether two combinations share dimensions.
Rows
12,000
Domain
physics
Format
data.parquet, one row per example
Licence
cc-by-4.0
Built for
supervised fine-tuning (SFT) experiments on Adaption AutoScientist
Columns
Column
Description… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-si-dimensional-analysis-12k.VisOnlyQA_eval_analysis_5
VisOnlyQA
🌐 Project Website | 📄 Paper | 🤗 Dataset | 🔥 VLMEvalKit
This repository contains the code and data for the paper "VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information".
VisOnlyQA is designed to evaluate the visual perception capability of large vision language models (LVLMs) on geometric information of scientific figures. The evaluation set includes 1,200 mlutiple choice questions in 12 visual perception tasks on 4… See the full description on the dataset page: https://huggingface.co/datasets/ryokamoi/VisOnlyQA_eval_analysis_5.Adversarial-Agent-Intent-Safety-Analysis-240K
Adversarial Agent Intent Safety Analysis 240K
Abstract
The Adversarial-Agent-Intent-Safety-Analysis-240K is a deterministically structured dataset featuring 242,454 context-rich adversarial prompts and safety evaluations. Engineered strictly for training frontier command-and-control models, guardrail classifiers, and red-teaming agents, it encourages models to parse multi-layered intention across 126 critical risk vectors.
This design trains models to decouple the surface… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Adversarial-Agent-Intent-Safety-Analysis-240K.sefd-archive-100k-analysis-sample-qwen3-20260524
SEFD Archive 100k Analysis Sample Qwen3 20260524
Retained artifacts for the completed archive-wide 100,000-filing Stanford EDGAR Filings Dataset (SEFD) analysis sample used in the arXiv paper update. The sample contains 2,971,490,909 final SEFD tokens, counted with the Qwen3-1.7B tokenizer.
This repository is a new versioned artifact and intentionally does not replace the earlier sfd-archive-100k-analysis-sample repository used for the original conference submission.
Included:… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/sefd-archive-100k-analysis-sample-qwen3-20260524.VisOnlyQA_eval_analysis_3
VisOnlyQA
🌐 Project Website | 📄 Paper | 🤗 Dataset | 🔥 VLMEvalKit
This repository contains the code and data for the paper "VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information".
VisOnlyQA is designed to evaluate the visual perception capability of large vision language models (LVLMs) on geometric information of scientific figures. The evaluation set includes 1,200 mlutiple choice questions in 12 visual perception tasks on 4… See the full description on the dataset page: https://huggingface.co/datasets/ryokamoi/VisOnlyQA_eval_analysis_3.RAG-Chunk-Analysis
Description
The datasets contain human evaluation of retrieved chunks from agriculture documents for actual user queries.
Each chunk is marked as relevant and irrelevant. The relevant and irrelevant portion of the chunks are mentioned in a separate columns.
The dataset consists of multiple XLS files and each XLS file has multiple sheets corresponding to the content for the value chain.
The queries are taken from the actual user questions onf farmer.chat prototype bots.
For each user… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/RAG-Chunk-Analysis.sales-analysisVisOnlyQA_eval_analysis_4
VisOnlyQA
🌐 Project Website | 📄 Paper | 🤗 Dataset | 🔥 VLMEvalKit
This repository contains the code and data for the paper "VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information".
VisOnlyQA is designed to evaluate the visual perception capability of large vision language models (LVLMs) on geometric information of scientific figures. The evaluation set includes 1,200 mlutiple choice questions in 12 visual perception tasks on 4… See the full description on the dataset page: https://huggingface.co/datasets/ryokamoi/VisOnlyQA_eval_analysis_4.Circuit-Analysis-Reasoning-Sample
⚡ EngineeringWays Data Lab: Circuit Analysis Reasoning Dataset (Free Sample)
This is a free 50-item sample of the EngineeringWays Circuit Analysis Reasoning Dataset. It is designed specifically for fine-tuning Large Language Models (LLMs) in advanced STEM problem-solving, featuring strict Chain-of-Thought (CoT) reasoning.
Want the complete, deduplicated 592-item master dataset? 👉 Get the LoRA-Ready Master File on Payhip
🚀 Dataset Overview
Most math and physics… See the full description on the dataset page: https://huggingface.co/datasets/EngineeringWays/Circuit-Analysis-Reasoning-Sample.cve-analysis
CVE & Vulnerability Analysis Dataset
A comprehensive vulnerability analysis and CVE research dataset. Each row is a detailed security analysis covering root cause, exploitation methodology, detection rules (Sigma/Splunk/Suricata), CVSS v3.1 scoring, MITRE ATT&CK mapping, and remediation guidance — verified by the same model in an independent review pass.
Overview
This dataset contains 9,999 structured vulnerability analyses across 20 security domains. Unlike simple… See the full description on the dataset page: https://huggingface.co/datasets/sh111111111111111/cve-analysis.VisOnlyQA_eval_analysis_2
VisOnlyQA
🌐 Project Website | 📄 Paper | 🤗 Dataset | 🔥 VLMEvalKit
This repository contains the code and data for the paper "VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information".
VisOnlyQA is designed to evaluate the visual perception capability of large vision language models (LVLMs) on geometric information of scientific figures. The evaluation set includes 1,200 mlutiple choice questions in 12 visual perception tasks on 4… See the full description on the dataset page: https://huggingface.co/datasets/ryokamoi/VisOnlyQA_eval_analysis_2.sfd-archive-100k-analysis-sample
SFD Archive 100k Analysis Sample
Retained artifacts for the completed archive-wide 100k-filing SFD analysis sample. The run processed 99,895 filings, produced 99,607 successful parses, and contains 2,698,335,638 final SFD tokens. This repository contains aggregate run metadata, processed accessions, and paper-analysis CSV/JSON metrics. Per-filing markdown outputs and temporary raw SEC downloads are not included.
Customer_Behavior_Analysis
Customer Behavior Analysis & CRM Intelligence Dataset (Thailand Context)
A specialized instruction-tuning dataset designed to train Large Language Models (LLMs) to serve as an Internal CRM Intelligence Copilot tailored specifically for the Thai market and consumer landscape.
The dataset bridges quantitative transaction logs (RFM, usage telemetry) with qualitative customer psychological theories within the local Thai business ecosystem (e.g., LINE OA interactions… See the full description on the dataset page: https://huggingface.co/datasets/yothinS/Customer_Behavior_Analysis.RAG-Chunk-Analysis
Description
The datasets contain human evaluation of retrieved chunks from agriculture documents for actual user queries.
Each chunk is marked as relevant and irrelevant. The relevant and irrelevant portion of the chunks are mentioned in a separate columns.
The dataset consists of multiple XLS files and each XLS file has multiple sheets corresponding to the content for the value chain.
The queries are taken from the actual user questions onf farmer.chat prototype bots.
For each user… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/RAG-Chunk-Analysis.VAB-vulnerability-analysis-benchmark
FBE and VAB
Two small benchmarks for security code analysis. Both grade without an LLM judge, so runs are cheap
and repeatable.
FBE (find-the-bug)
14 code snippets, each with one planted vulnerability. Ask the model to analyze the code, then check
whether it actually found the flaw.
Grading uses concept groups: the answer has to contain at least one synonym from every required group.
Four numbers come out:
found, did it identify the real vulnerability (this is… See the full description on the dataset page: https://huggingface.co/datasets/MK4-Research/VAB-vulnerability-analysis-benchmark.RAG-Chunk-Analysis
Description
The datasets contain human evaluation of retrieved chunks from agriculture documents for actual user queries.
Each chunk is marked as relevant and irrelevant. The relevant and irrelevant portion of the chunks are mentioned in a separate columns.
The dataset consists of multiple XLS files and each XLS file has multiple sheets corresponding to the content for the value chain.
The queries are taken from the actual user questions onf farmer.chat prototype bots.
For each… See the full description on the dataset page: https://huggingface.co/datasets/Keerthana092005/RAG-Chunk-Analysis.RAG-Chunk-Analysis
Description
The datasets contain human evaluation of retrieved chunks from agriculture documents for actual user queries.
Each chunk is marked as relevant and irrelevant. The relevant and irrelevant portion of the chunks are mentioned in a separate columns.
The dataset consists of multiple XLS files and each XLS file has multiple sheets corresponding to the content for the value chain.
The queries are taken from the actual user questions onf farmer.chat prototype bots.
For each… See the full description on the dataset page: https://huggingface.co/datasets/deepak-1011/RAG-Chunk-Analysis.market-analysis-africa
Market Analysis & News — African Context Dataset
Instruction-tuning dataset covering African market analysis and business news: stock exchanges (USE, NSE, NGX, JSE), commodity markets (coffee, cocoa, gold, oil), regional trade (EAC, AfCFTA, ECOWAS), macroeconomic indicators, startup/VC ecosystems, real estate, and sector intelligence — grounded via web search, generated with gemini-2.5-flash.
Dataset Details
Rows: 187
Regions covered: Uganda, Kenya, Tanzania… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/market-analysis-africa.sales_analysis_queriessales_analysis1investment_analysis
코스피 상장 기업 공시정보 기반 투자 리포트 데이터셋
이 데이터셋은 국내 코스피 상장 기업의 공시정보를 바탕으로, 투자 전문가들이 활용할 수 있는 심층적 분석과 투자 전략 제안을 목표로 제작되었습니다. 특히, 이 데이터셋은 GPT 파인튜닝에 최적화된 구조로 설계되어 있어, 다양한 역할(role)을 포함한 메시지 기반의 대화 형식으로 구성되어 있습니다.
데이터셋 구조
데이터셋은 JSONL 포맷으로 제공되며, 각 항목은 GPT 파인튜닝에 최적화된 메시지 형식을 따릅니다. 주요 구성은 다음과 같습니다:
messages: 메시지 배열 형태로 구성되어 있으며, 각 메시지는 아래와 같은 역할을 가집니다.
system: 모델의 역할과 행동 지침을 정의합니다.예시: "당신은 기업 재무 및 투자 분석 전문가입니다. 참고 컨텍스트를 기반으로 사용자 질문에 대해 정확하고 논리적으로 답변하세요."
user: 사용자의 질문과 컨텍스트(예시 데이터, 재무제표… See the full description on the dataset page: https://huggingface.co/datasets/MLOpsEngineer/investment_analysis.llama-security-log-analysis
LLaMA Security Log Analysis (Clean Format)
A security log analysis dataset converted from mkenfenheuer/llama-security-llm with all LLaMA special tokens removed for clean GPT/ShareGPT format compatibility.
Dataset Description
This dataset contains 4,189 examples of security log analysis conversations. The original dataset had LLaMA 3 formatting tokens (<|begin_of_text|>, <|start_header_id|>, etc.) which have been cleanly removed to create a universal conversation format.… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/llama-security-log-analysis.pt-it-analysisfinancial_analysismedqa-phi4-failure-analysisThis dataset contains a comprehensive log of reasoning and answers generated by microsoft/Phi-4-mini-instruct, evaluated on medalpaca/medical_meadow_medqa (USMLE) dataset.
This dataset represents instances where model got the answer right as well as wrong. All examples includes reasoning. The inference was performed locally on Macbook (M-series) using the MLX-LM framework (The model parameters were: temp: 0.3, max_tokens: 300).
