datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
adaption-financial-reasoning-steps
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-financial_reasoning_steps
This dataset contains pairs of financial analysis questions and their corresponding step-by-step reasoning processes to derive numerical answers. Each entry includes a specific query about corporate metrics like growth rates, percentages, or net changes, followed by explicit arithmetic operations and a final calculated value. The content is structured to… See the full description on the dataset page: https://huggingface.co/datasets/asadullahdogarr/adaption-financial-reasoning-steps.adaption-financial-qa-pairs
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-financial_qa_pairs
This dataset consists of question-and-answer pairs focused on extracting specific financial metrics from corporate reports. The prompts inquire about percentages, monetary values, growth rates, and operational statistics such as store counts or debt changes. Each completion provides a precise numerical answer derived from financial statements or related… See the full description on the dataset page: https://huggingface.co/datasets/dipanjann/adaption-financial-qa-pairs.adaption-financial-tat-qa-pairs
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-financial_tat_qa_pairs
This dataset consists of question-and-answer pairs derived from corporate financial reports, specifically Form 10-K filings. The prompts inquire about specific numerical metrics, year-over-year comparisons, and qualitative explanations for financial trends or organizational changes. Completions provide precise values, percentages, or direct textual excerpts… See the full description on the dataset page: https://huggingface.co/datasets/dipanjann/adaption-financial-tat-qa-pairs.adaption-defi-wallet-risk-classification
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-defi_wallet_risk_classification
This dataset contains prompt-completion pairs for classifying the 14-day risk outcomes of DeFi wallets on various EVM networks. Each entry provides behavioral features such as transaction counts, action ratios, and concentration metrics within a specific feature window to predict a binary risk label. The completions offer a concise justification for… See the full description on the dataset page: https://huggingface.co/datasets/samscript18/adaption-defi-wallet-risk-classification.adaption-python-mental-execution-traces
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-python_mental_execution_traces
This dataset contains pairs of Python 3 code snippets and their corresponding mental execution results, including exact standard output and concise variable traces. Each sample challenges the model to simulate code logic involving lists, matrices, loops, and conditional statements without actual execution. The completions provide both the final printed… See the full description on the dataset page: https://huggingface.co/datasets/ILoveBuns/adaption-python-mental-execution-traces.adaption-financial-table-qa
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-financial_table_qa
This dataset contains question-answer pairs derived from corporate financial statements, including balance sheets and income statements for various companies. The prompts require extracting specific line items, calculating year-over-year percentage changes, or converting scaled values (millions/thousands) to absolute dollar amounts. Each completion provides the… See the full description on the dataset page: https://huggingface.co/datasets/RaffiArdhi/adaption-financial-table-qa.adaption-india-medical-triage-safety
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-india_medical_triage_safety
This dataset contains prompt-completion pairs for medical triage scenarios specific to India, covering emergencies like seizures, snake bites, and chest pain across various Indian languages. Each entry classifies severity, provides safe response guidance, lists unsafe actions to avoid, and specifies escalation steps such as calling emergency services. The… See the full description on the dataset page: https://huggingface.co/datasets/jadhavmanasi70/adaption-india-medical-triage-safety.adaption-sec-financial-arithmetic-dataset
SEC Financial Arithmetic Dataset — Adaption AutoScientist Challenge
Powered by Adaptive Data — Adaption Labs
What This Dataset Teaches
This dataset trains a model to extract numbers from SEC filing tables and execute verified multi-step arithmetic — every answer is cross-checked against a gold reasoning program:
Task
Source
Example
Table Variable Extraction
FinQA
"From this 10-K table, extract 2021 and 2022 revenue values"
Multi-Step Arithmetic… See the full description on the dataset page: https://huggingface.co/datasets/narendarcodes/adaption-sec-financial-arithmetic-dataset.adaption-urdu-edu-cultural-reasoning
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-urdu_edu_cultural_reasoning
This dataset contains a mixed collection of question-answer pairs and linguistic tasks presented in both English and Urdu. The content spans multiple domains including history, biology, geography, and Urdu literature, featuring multiple-choice questions, translation exercises, and poetic composition prompts. Samples include historical treaty analysis… See the full description on the dataset page: https://huggingface.co/datasets/abdullah693/adaption-urdu-edu-cultural-reasoning.adaption-vi-medical-safety-jailbreak
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-vi_medical_safety_jailbreak
This dataset contains Vietnamese prompts and completions focused on pharmaceutical safety and medical ethics scenarios. The samples present conflicting instructions where users attempt to bypass safety protocols, such as dispensing medication despite severe allergic reactions or falsifying patient status. The corresponding completions demonstrate the… See the full description on the dataset page: https://huggingface.co/datasets/melanieyes/adaption-vi-medical-safety-jailbreak.adaption-louisville-data-center-docs-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-louisville_data_center_docs (augmented)
This dataset contains planning commission staff reports, zoning code excerpts, and news transcripts regarding hyperscale data center developments in Louisville, Kentucky. The documents detail specific project proposals, such as the Camp Ground Road facility, including technical reviews on traffic, water usage, and environmental impact.… See the full description on the dataset page: https://huggingface.co/datasets/JaySmith502/adaption-louisville-data-center-docs-augmented.adaption-kenyan-finance-reasoning-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-kenyan_finance_reasoning (augmented)
This dataset contains prompt-completion pairs focused on personal finance calculations and strategic reasoning within the Kenyan economic context, covering topics like Chama rotations, KRA tax deductions, and debt repayment strategies. Each sample provides step-by-step mathematical derivations and comparative analyses to guide financial… See the full description on the dataset page: https://huggingface.co/datasets/Gro97/adaption-kenyan-finance-reasoning-augmented.adaption-piguard-translation-handoff
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-piguard_translation_handoff
This dataset contains 100 English prompts curated for translation guardrail research, comprising a balanced mix of 50 benign instructions and 50 prompt injection attempts. The samples include diverse content such as jailbreak personas, requests for harmful actions like spyware installation, and complex instruction overrides. Each entry is… See the full description on the dataset page: https://huggingface.co/datasets/melanieyes/adaption-piguard-translation-handoff.adaption-math-logic-verification-pairs-augmented-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-math_logic_verification_pairs (augmented) (augmented)
This dataset contains prompt-completion pairs focused on verifying mathematical and logical claims across algebra, calculus, logic, and topology. Each prompt presents a problem statement with a claimed solution, while the completion provides a step-by-step verification determining validity and correcting errors where necessary.… See the full description on the dataset page: https://huggingface.co/datasets/Gro97/adaption-math-logic-verification-pairs-augmented-augmented.adaption-med-safety-preference-pairs
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-med_safety_preference_pairs
This dataset consists of medical preference pairs comparing chosen and rejected model completions across symptom evaluation, diagnosis, medications, and pharmacology. Scenarios include ambiguous patient presentations requiring clarifying questions, as well as critical safety checks regarding drug interactions, contraindications, and proper dosages. All… See the full description on the dataset page: https://huggingface.co/datasets/Gro97/adaption-med-safety-preference-pairs.adaption-math-logic-verification-pairs-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-math_logic_verification_pairs (augmented)
This dataset contains prompt-completion pairs focused on verifying mathematical and logical claims across algebra, calculus, logic, and topology. Each prompt presents a problem statement with a claimed solution, while the completion provides a step-by-step verification determining validity and correcting errors where necessary. The content… See the full description on the dataset page: https://huggingface.co/datasets/Gro97/adaption-math-logic-verification-pairs-augmented.adaption-amharic-text-corpus
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-amharic_text_corpus
This dataset comprises over 700,000 Amharic text documents formatted as line-delimited JSON, covering diverse topics such as history, religion, politics, and product descriptions. Each entry contains a single string field with native Amharic content, including some samples with mixed languages or placeholder values. It is designed for text… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/adaption-amharic-text-corpus.adaption-agent-memory-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-agent-memory (augmented)
This dataset contains samples of conversations between a user and an assistant, paired with the corresponding structured memory updates extracted for long-term storage. Each sample includes the full dialogue context, existing memory state, and the resulting JSON output containing new narrative summaries and atomic facts with specific keys and values.… See the full description on the dataset page: https://huggingface.co/datasets/huyxdang/adaption-agent-memory-augmented.adaption-hinglish-transliterate-dataset
Adaption Hinglish Transliterate Dataset
Dataset Description
This dataset contains 77,471 pairs of raw Hindi text captured via Automatic Speech Recognition (ASR) in Devanagari script and their corresponding clean transliterations into Romanized Hinglish. The samples demonstrate the correction of ASR artifacts and the application of Anglicized Hinglish conventions while preserving the original meaning.
Each entry consists of an original system prompt instructing… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/adaption-hinglish-transliterate-dataset.adaption-moroccan-darija-prompts-trilingual-codeswitch-chat-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-moroccan_darija_prompts & trilingual_codeswitch_chat (augmented)
This dataset consists of short conversational prompts written in Moroccan Darija, covering topics like shopping, social interactions, and daily inquiries. Each entry contains a single prompt with a null completion, indicating it is likely intended for instruction tuning or completion generation tasks. The content… See the full description on the dataset page: https://huggingface.co/datasets/oumayma03/adaption-moroccan-darija-prompts-trilingual-codeswitch-chat-augmented.adaption-african-history-discoveries
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-african_history_discoveries
This dataset consists of instruction-response pairs covering contemporary discoveries and reassessments in African history from 2020 to 2026. Samples feature news snippets and research summaries alongside factual contextual analyses of archaeological finds, oral tradition documentations, genetic studies, and colonial-era historical re-evaluations. Each… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/adaption-african-history-discoveries.adaption-agriintel-contextual-yield-reasoning-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-AgriIntel-Contextual-Yield-Reasoning-v1
A deeply evolved, instruction-tuned agricultural dataset engineered for precision agronomy modeling. This dataset bridges the gap between raw regional telemetry and generative AI reasoning by synthesizing historical crop performance with static soil chemistry metrics (N, P, K, pH) and historical macro-climatic atmospheric data (annual rainfall… See the full description on the dataset page: https://huggingface.co/datasets/Shravanthmvqwerty/adaption-agriintel-contextual-yield-reasoning-v1.adaption-math-logic-verification-pairs
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-math_logic_verification_pairs
This dataset contains prompt-completion pairs focused on verifying mathematical and logical claims across algebra, calculus, logic, and topology. Each prompt presents a problem statement with a claimed solution, while the completion provides a step-by-step verification determining validity and correcting errors where necessary. The content covers… See the full description on the dataset page: https://huggingface.co/datasets/Gro97/adaption-math-logic-verification-pairs.adaption-financial-math-qa
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-financial_math_qa
This dataset contains question-and-answer pairs focused on personal finance calculations, including compound interest, loan amortization, tax brackets, and retirement planning. Each sample provides a specific financial scenario in the prompt and a detailed, step-by-step mathematical derivation in the completion. The responses explain the underlying formulas… See the full description on the dataset page: https://huggingface.co/datasets/uditjain/adaption-financial-math-qa.adaption-formal-step-verification-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-formal_step_verification (augmented)
This dataset contains pairs of prompts and completions focused on verifying the validity of logical arguments, algebraic derivations, calculus operations, and combinatorial steps. Each entry requires the model to determine if a claimed conclusion or transformation is correct, often providing counterexamples or identifying specific error indices… See the full description on the dataset page: https://huggingface.co/datasets/Gro97/adaption-formal-step-verification-augmented.adaption-agent-memory-augmented-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-agent-memory (augmented)
This dataset contains samples of conversations between a user and an assistant, paired with the corresponding structured memory updates extracted for long-term storage. Each sample includes the full dialogue context, existing memory state, and the resulting JSON output containing new narrative summaries and atomic facts with specific keys and values.… See the full description on the dataset page: https://huggingface.co/datasets/huyxdang/adaption-agent-memory-augmented-v1.adaption-recipe-ingredient-validation
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-recipe_ingredient_validation
This dataset consists of prompt-completion pairs designed to test ingredient relevance for specific recipes. Each entry presents a recipe title and a candidate ingredient, requiring a binary classification of whether the ingredient belongs in the dish. The completions provide the correct label as either 'BELONGS' or 'NOT_BELONGS' based on culinary logic.… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-recipe-ingredient-validation.adaption-hr-advisory-onet
HR Advisory Instruction Dataset (O*NET-grounded)
Instruction-tuning data for HR advisory work — job design, hiring, assessment, internal mobility, workforce analytics and tooling — with every factual claim traceable to a named O*NET occupation record.
Built for the Adaption Labs AutoScientist Challenge Part 2, HR track.
What is in it
Rows
5,415 (4,836 train / 579 eval)
Task families
19
Occupations covered
907 of 923 available
Response length… See the full description on the dataset page: https://huggingface.co/datasets/miscusi/adaption-hr-advisory-onet.adaption-market-analysis-sec
Market Analysis & News Instruction Dataset (SEC XBRL-grounded)
Instruction-tuning data for financial analysis — fundamentals, growth and ratio arithmetic, trend and risk reading, filing navigation and comparability caveats — built from real XBRL facts, with every stated figure independently re-derived.
Built for the Adaption Labs AutoScientist Challenge Part 2, Market Analysis & News track.
What is in it
Rows
5,068 (4,501 train / 567 eval)
Task… See the full description on the dataset page: https://huggingface.co/datasets/miscusi/adaption-market-analysis-sec.adaption-agronomy-qa-pairs
East Africa Agronomy QA
Agricultural AI for the languages and people feeding East Africa
East Africa's food system is enormous. Agriculture remains at the centre of the region's economy and livelihoods, while Africa as a whole has around 300 million people employed in agrifood systems and the highest regional share of employment in agrifood systems globally, at 64.5%. Agriculture accounts for roughly 74.4% of Africa's agrifood-system employment.… See the full description on the dataset page: https://huggingface.co/datasets/RayNene/adaption-agronomy-qa-pairs.
