Team Ai
Datasetpublic

AmanPriyanshu/RLVR-Env-Retrieval-Source-code-search-net-python

RLVR-Env-Retrieval-Source-code-search-net-python RLVR-ready retrieval environment derived from Nan-Do/code-search-net-python. Author: Aman Priyanshu What Is This A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches through… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-code-search-net-python.

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
0likes22downloads
README.md76 linesDownload Raw Back to root
1---2license: apache-2.03task_categories:4  - text-retrieval5  - question-answering6language:7  - en8tags:9  - retrieval10  - rlvr11  - search12  - distractor-mining13size_categories:14  - 100K<n<1M15---16 17# RLVR-Env-Retrieval-Source-code-search-net-python18 19RLVR-ready retrieval environment derived from [Nan-Do/code-search-net-python](https://huggingface.co/datasets/Nan-Do/code-search-net-python).20 21**Author:** [Aman Priyanshu](https://huggingface.co/AmanPriyanshu)22 23## What Is This24 25A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches through distractors to find the correct chunk(s).26 27**Domain:** Python open-source functions from GitHub (CodeSearchNet)28 29## Source30 31Derived from [Nan-Do/code-search-net-python](https://huggingface.co/datasets/Nan-Do/code-search-net-python) (455,243 unique functions).32Original license: **Apache 2.0** — retained here.33 34## Schema35 36### qa.parquet (100,000 rows)37 38| Column | Type | Description |39|---|---|---|40| `qa_id` | string | Unique ID (`search_py_0`, `search_py_1`, ...) |41| `question` | string | The retrieval query |42| `gt_chunks` | JSON string | List of ground-truth chunk texts. 1 target code chunk per question (the function matching the summary) |43| `random_chunks` | JSON string | List of random distractor texts. ~500 random code chunks (>=20 chars, deduplicated against gt and similar) |44| `similar_chunks` | JSON string | List of hard-negative distractor texts. ~178 similar chunks via MiniLM cosine (<0.97) + char trigram edit-distance (<0.97 seq ratio), deduplicated |45 46### metadata.parquet (100,000 rows)47 48| Column | Type | Description |49|---|---|---|50| `qa_id` | string | Matches qa.parquet |51| ... | ... | chunk_idx, func_name, repo, char_count |52 53### chunks.parquet54 55455,243 code chunks with MiniLM embeddings. Kept for reference — not needed at inference time.56 57## Deduplication58 59Within each row: gt > similar > random priority. No chunk text appears in more than one column per row. Similar chunks are internally deduplicated. Random chunks are filtered against both gt and similar.60 61## How To Use62 63```python64import json65import pyarrow.parquet as pq66 67t = pq.read_table("qa.parquet")68row = {col: t.column(col)[0].as_py() for col in t.column_names}69gt = json.loads(row["gt_chunks"])70distractors = json.loads(row["random_chunks"]) + json.loads(row["similar_chunks"])71```72 73## License74 75Apache 2.0 (inherited from source dataset).76