datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Scientific-Code-and-Analysis-QA
RegalFire Scientific code and analysis QA
RegalFire — AI Data Foundry
Scientific forum QA containing mechanically extracted preformatted code. Code is retained exactly after HTML entity decoding; no execution is claimed.
Verified scope
Records: 72; distinct source threads: 72; unique answers represented: 99.
Domain thread counts: {"statistics": 34, "computational_science": 37, "biology": 1}.
Actual record splits: {"train": 48, "holdout": 12, "test": 6… See the full description on the dataset page: https://huggingface.co/datasets/RegalFire/Scientific-Code-and-Analysis-QA.Code-Mixed-Sentiment-Analysis-Dataset
Dataset Generation:
Initially, we select the Amazon Review Dataset as our base data, referenced from Ni et al. (2019)[^1]. We randomly extract 100,000 instances from this dataset. The original labels in this dataset are ratings, scaled from 1 to 5. For our specific task, we categorize them into Positive (rating > 3), Neutral (rating = 3), and Negative (rating < 3), ensuring a balanced number of instances for each label. To generate the synthetic Code-mixed dataset, we apply two… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Sentiment-Analysis-Dataset.code_completion_for_data_analysiscode-analysis-sft-qwen38-v2instruct_code_for_data_analysiscode-gen-lcb-error-analysispedagogia-c-code-analysis-sftmmlu-olmo37b-stage2-math-code-analysis-baselinemmlu-olmo37b-stage2-math-code-analysis-with-context
