Team Ai
20 results

codeswitching

liva-ai /code-switching-asrIf the samples in the dataset viewer don't load, they can also be accessed here. Code-Switching ASR Code-switching ASR is a speech dataset of code-switching between English and medium to low resource languages. Samples include en-sw, en-pcm, en-yo, and en-tl but we can deliver any of the languages listed here: https://huggingface.co/datasets/liva-ai/yapdo-convo (the hours have yet to be updated as of 07/14/2026 as we have much higher volume now - and we can easily collect more… See the full description on the dataset page: https://huggingface.co/datasets/liva-ai/code-switching-asr.audioautomatic-speech-recognitionn<1K0 likes200 downloads3mo agoHugging FaceMalikeh1375 /code-switching-tokenizer-robustness Code-Switching Dataset for Tokenizer Robustness Analysis Dataset Description This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models. Purpose Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.texttext-generation1K<n<10K2 likes129 downloads1y agoHugging Facecode-switching /question-answertext1K<n<10K0 likes82 downloads1mo agoHugging Facecode-switching /text-summarizationtextsummarizationn<1K0 likes78 downloads1mo agoHugging Facedevrahulbanjara /ne-en-codeswitching-asr-technical-interview Dataset Summary This dataset contains audio recordings and text transcripts of Nepali-English code-switched speech in the context of technical interviews. It is specifically designed to handle the linguistic complexities of Nepali software engineers, developers, and IT professionals who frequently mix English technical terminology (e.g., AWS, S3 lifecycle policies, RAG pipelines, VPC peering) with conversational Nepali grammar. It is an excellent resource for fine-tuning ASR models… See the full description on the dataset page: https://huggingface.co/datasets/devrahulbanjara/ne-en-codeswitching-asr-technical-interview.audioautomatic-speech-recognitionn<1K3 likes61 downloads7mo agoHugging Faceanishka /CodeSwitching-TE-ENAnCora Catalan NER. This is a dataset for Named Eentity Reacognition (NER) from Ancora corpus adapted for Machine Learning and Language Model evaluation purposes. Since multiwords (including Named Entites) in the original Ancora corpus are aggregated as a single lexical item using underscores (e.g. "Ajuntament_de_Barcelona") we splitted them to align with word-per-line format, and added conventional Begin-Inside-Outside (IOB) tags to mark and classify Named Entites. We did not filter out the different categories of NEs from Ancora (weak and strong). We did 6 minor edits by hand. AnCora corpus is used under [CC-by] (https://creativecommons.org/licenses/by/4.0/) licence. This dataset was developed by BSC TeMU as part of the AINA project, and to enrich the Catalan Language Understanding Benchmark (CLUB).token-classification2 likes39 downloads3y agoHugging Face