semaj83/ctmatch_classification
CTMatch Classification Dataset This is a combined set of 2 labelled datasets of: topic (patient descriptions), doc (clinical trials documents - selected fields), and label ({0, 1, 2}) triples, in jsonl format. (Somewhat of a duplication of some of the ir_dataset also available on HF.) These have been processed using ctproc, and in this state can be used by various tokenizers for fine-tuning (see ctmatch for examples). These 2 datasets contain no patient identifying information are openly… See the full description on the dataset page: https://huggingface.co/datasets/semaj83/ctmatch_classification.
Delete test_index2docid.txt
Delete test_doc_texts.txt
Delete test_doc_embeddings.txt
Delete test_doc_categories.txt
Delete index2docid.txt
Delete doc_texts.txt
Delete doc_embeddings.txt
Delete doc_categories.txt
Update README.md
Rename test_index2id.txt to test_index2docid.txt
Rename index2id.txt to index2docid.txt
Upload 2 files
Delete index2docid.txt
Delete test_index2id.txt
Rename test_categories.txt to test_doc_categories.txt
Rename test_embeddings.txt to test_doc_embeddings.txt
Upload test_doc_texts.txt
Upload 3 files
Rename index2docid.csv to index2docid.txt
Rename doc_embeddings.csv to doc_embeddings.txt
Rename doc_categories.csv to doc_categories.txt
Upload 2 files
Delete doc_embeddings.csv
Delete doc_categories.csv
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Upload doc_texts.txt
Rename index2id.csv to index2docid.csv
Upload 3 files
Delete categories
Delete doc_embeddings.csv
Delete index2id.csv
Delete doc_categories.csv
Upload 2 files
Delete idx2id.csv
Upload 2 files
Update README.md
Upload combined_classifier_data.jsonl
Delete trec_data.jsonl
Delete kz_data.jsonl
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
