Team Ai
Datasetpublic

projecte-aina/CA-EU_Parallel_Corpus

Dataset Card for CA-EU Parallel Corpus Dataset Summary The CA-EU Parallel Corpus is a Catalan-Basque synthetic dataset of parallel sentences created to support the use of co-official languages from Spain, such as Catalan and Basque, in NLP tasks, specifically Machine Translation. Supported Tasks and Leaderboards The dataset can be used to train Bilingual Machine Translation models between Basque and Catalan in any direction, as well as… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CA-EU_Parallel_Corpus.

sourceHugging Facecc-by-nc-sa-4.0updated 2y agoView on Hugging Face
0likes44downloads
23 commits on main
f02bc332y ago

Update README.md

mmarimon
eac11602y ago

Add parquet file description

fdelucaf
c3535662y ago

Upload parquet

fdelucaf
19bd0c72y ago

Delete train_clean.eu

fdelucaf
88580ca2y ago

Minor changes to datacard

fdelucaf
a68740e2y ago

Upload ca-eu_FULL.eu

AudreyVM
b0467902y ago

update dataset card

AudreyVM
0ac1cb22y ago

Delete train_clean.ca

AudreyVM
f4225bf2y ago

fix presence of unknowns in corpus

AudreyVM
b3110603y ago

Update README.md

fdelucaf
27ea2fc3y ago

Update README.md

fdelucaf
5fe05893y ago

Update README.md

fdelucaf
8294ca43y ago

Update README.md

AudreyVM
e4c42d03y ago

Update README.md

fdelucaf
6c4cd613y ago

Update README.md

fdelucaf
cf89de03y ago

Upload 2 files

fdelucaf
c1d5bfd3y ago

Update README.md

fdelucaf
7069b8b3y ago

Update README.md

fdelucaf
a8777233y ago

Update README.md

fdelucaf
f6f18bb3y ago

Update README.md

fdelucaf
aa4d0363y ago

Update README.md

fdelucaf
d7f644a3y ago

Create README.md

fdelucaf
90adc893y ago

initial commit

fdelucaf