datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
waifuc_unittestUnittest files for deepghs/waifuc.
UnitTests
UnitTests
Task description
Evaluation of unit-test generation for functions and methods in five programming languages (Java, Python, Go, JavaScript, and C#). Dataset contains 2500 tasks.
Evaluated skills: Instruction Following, Long Context Comprehension, Synthesis, Testing
Contributors: Alena Pestova, Valentin Malykh
Motivation
Unit testing is an important software development practice in which individual components of a software system are evaluated in… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/UnitTests.CodeRM-UnitTest
Dataset Description
CodeRM-UnitTest dataset originates from the paper: Dynamic Scaling of Unit Tests for Code Reward Modeling available on arXiv.
You can visit the homepage to learn more about the paper.
It is a curated collection of high-quality synthetic Python unit tests, derived from two prominent code instruction tuning
datasets: CodeFeedback-Filtered-Instruction and the training
set of TACO. This dataset is used for training
CodeRM-8B, a small yet powerful unit test… See the full description on the dataset page: https://huggingface.co/datasets/KAKA22/CodeRM-UnitTest.UnitTest-CleanedGO-UNITTEST-BENCHCPP-UNITTEST-BENCH
Dataset Card for Open Source Code and Unit Tests
Dataset Details
Dataset Description
This dataset contains c++ code snippets and their corresponding ground truth unit tests collected from various open-source GitHub repositories. The primary purpose of this dataset is to aid in the development and evaluation of automated testing tools, code quality analysis, and LLM models for test generation.
Curated by: Vaishnavi Bhargava
Language(s): C++… See the full description on the dataset page: https://huggingface.co/datasets/Nutanix/CPP-UNITTEST-BENCH.UnitTestsPublicSWESwiss-SFT-Unittest-1K
Overview
SFT dataset for training SWE-Swiss models on the unit test generation task. The prompts contain issues sourced from SWE-Gym and SWE-smith, while the responses are generated by DeepSeek-R1-0528. To ensure quality, we filter out data where the generated unit tests do not perform as expected. A generated test is kept only if its execution results correctly distinguish between a set of correct and incorrect patches, mirroring the behavior of the repository's own test suite.… See the full description on the dataset page: https://huggingface.co/datasets/SWE-Swiss/SWESwiss-SFT-Unittest-1K.unit-test-v2
Dataset Card for "unit-test-v2"
More Information needed
python-unittest-fixed
python-unittest-fixed
A repaired and augmented version of rd320uetvnu/python-unittest.
Why this version exists
The original dataset had a column misalignment: in ~99% of records, the code
field defined one function while unittest_gen (and thinking) referenced a
different function. The (code, unit_test_org) columns and the
(unittest_gen, thinking, cov->unittest) columns came from two independent sources
that had been joined row-wise without aligning content. As a… See the full description on the dataset page: https://huggingface.co/datasets/rd320uetvnu/python-unittest-fixed.soundutils_unittest
Unittest Assets for SoundUtils
Assests for unittest for deepghs/soundutils.
unit-test-mocking-dataset-enriched-v2unit_test_generation⚠️ Note: The dataset symprompt_supp.jsonl is not created by us. We only supplemented this dataset with additional branch-level metadata (e.g., has_branch, total_branches) to enable coverage testing.
This helps users keep their workflows clean when determining whether branches exist, simplifying branch coverage calculation.
It originates from the paper:
Code-Aware Prompting: A Study of Coverage Guided Test Generation in Regression Setting using LLM
— Gabriel Ryan, Siddhartha Jain, Mingyue… See the full description on the dataset page: https://huggingface.co/datasets/Code-TREAT/unit_test_generation.unit-test_PDFfolder
Required installation
pip3 install pypdf2 pdf2image
sudo apt-get install poppler-utils
UnitTest-Finetuning Dataset Card for UnitTest-Finetuning Corpus
Dataset Summary
UnitTest-Finetuning corpus is a dataset of 1.48 million instruction prompts for the development tasks in AD and TDD, across about 5.5K Java and 16.2K Python open-source projects on GitHub.
Supported Tasks
Method-Test Dataset: The dataset can be used to train a model for test completion and test generation tasks in Agile Development
Docstring-Test-Method Dataset: The dataset can be used to train a model for… See the full description on the dataset page: https://huggingface.co/datasets/Arain/UnitTest-Finetuning.unit-test-mocking-dataset-enriched-v3realutils_unittestunit-test-v1
Dataset Card for "unit-test-v1"
More Information needed
monkey_business_128_MATH_llama_70b_unittests_resultssafecode_unittestLog_MAS_UNITTESTunit-test-mocking-dataset-enrichedEmbedded_Unittest1
Dataset Card for "Embedded_Unittest1"
More Information needed
python-unittestunit-test-mocking-dataset-v3public_unittest_repoThis is an integration database of erai-raws, myanimelist and nyaasi. You can know which animes are the hottest ones currently, and which of them have well-seeded magnet links.
This database is refreshed daily.
Current Animes
5 animes, 45 episodes in total, Last updated on: 2024-07-21 01:08:35 CST.
ID
Post
Bangumi
Type
Episodes
Status
Score
Nyaasi
Magnets
Seeds
Downloads
Updated At
53802
2.5-jigen no RirisaTV
7 / 24
Currently Airing
7.29
Search
Download
32
900
2024-07-19… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/public_unittest_repo.unit-testingText_Normalization_Challenge_Unittests_Eng_Fraunit-test-mocking-dataset-v2Log_Filtered_MAS_UNITTEST_v1
