Team Ai
20 results

paraphrase

biglab /jitteredwebsites-merged-224-paraphrasedimage1M<n<10M1 likes418 downloads3y agoHugging Facemerionum /ru_paraphraser Dataset Card for ParaPhraser Dataset Summary ParaPhraser is a news headlines corpus annotated according to the following schema: 1: precise paraphrases 0: near paraphrases -1: non-paraphrases The Plus part is also available. It contains clusters of news headline paraphrases labeled automatically by a fine-tuned paraphrase detection BERT model.In order to load it: from datasets import load_dataset corpus = load_dataset('merionum/ru_paraphraser', data_files='plus.jsonl')… See the full description on the dataset page: https://huggingface.co/datasets/merionum/ru_paraphraser.texttext-classification1K<n<10K13 likes393 downloads4y agoHugging Facehumarin /chatgpt-paraphrasesThis is a dataset of paraphrases created by ChatGPT. Model based on this dataset is avaible: model We used this prompt to generate paraphrases Generate 5 similar paraphrases for this question, show it like a numbered list without commentaries: {text} This dataset is based on the Quora paraphrase question, texts from the SQUAD 2.0 and the CNN news dataset. We generated 5 paraphrases for each sample, totally this dataset has about 420k data rows. You can make 30 rows from a row from… See the full description on the dataset page: https://huggingface.co/datasets/humarin/chatgpt-paraphrases.text100K<n<1M61 likes342 downloads4y agoHugging FaceFaless /harvest_apples_with_agilex_piper_sim_ee_paraphrases20This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 25, "features": { "observation.state": { "dtype": "float32", "fps": 25, "shape": [ 8 ], "names": [ "ee.x", "ee.y", "ee.z", "ee.roll", "ee.pitch", "ee.yaw"… See the full description on the dataset page: https://huggingface.co/datasets/Faless/harvest_apples_with_agilex_piper_sim_ee_paraphrases20.tabularrobotics1M<n<10M0 likes305 downloads4mo agoHugging Facehpprc /paraphrase-qa日本語Wikipedia中のテキストを元に言い換えを生成し、その言い換えを元にクエリと回答をLLMに生成させたデータセットです。 出力にライセンス的な制約があるモデルを利用していないことと、元データとして日本語Wikipediaを利用していることから、CC-BY-SA 4.0ライセンスのもとでの配布とします。 textquestion-answering10M<n<100M2 likes200 downloads2y agoHugging FaceTurkuNLP /turku_paraphrase_corpusTurku Paraphrase Corpus is a dataset of 104,645 manually annotated Finnish paraphrases. The vast majority of the data is classified as a paraphrase either in the given context, or universally.text-classification100K<n<1M4 likes191 downloads2y agoHugging Face