Team Ai
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MrJackTung /cs-envi-dual-encoder-60audion<1K0 likes133 downloads5mo agoHugging Face02Unggi /modernbert_encoder_sp_seq_512_csedm_fold1tabularn<1K0 likes49 downloads2y agoHugging Face03lmms-lab-encoder /SciBench_Mathtextn<1K0 likes43 downloads1y agoHugging Face04Unggi /modernbert_encoder_sp_seq_512_dbe22kt_fold1tabularn<1K0 likes38 downloads2y agoHugging Face05HqH1111 /AutoMoT-PDM-Lite-BEV-Encoder-Indexes AutoMoT PDM-Lite BEV Encoder Indexes This dataset provides the prepared PDM-Lite JSONL indexes for AutoMoT training. Files pdm_lite_2hz_2tp_train_bev_encoder.jsonl pdm_lite_2hz_2tp_val_bev_encoder.jsonl Each row contains four historical front-camera paths in image, the current front-camera path in front, trajectory and route supervision, future-speed supervision, and a reference to the precomputed current-frame BEV feature: bev_encoder_feature… See the full description on the dataset page: https://huggingface.co/datasets/HqH1111/AutoMoT-PDM-Lite-BEV-Encoder-Indexes.texttext-generation100K<n<1M0 likes22 downloads3mo agoHugging Face06chungimungi /arxiv-hard-negatives-cross-encoderThis dataset contains hard negative examples generated using cross-encoders for training dense retrieval models. @misc{reimers2019sentencebertsentenceembeddingsusing, title={Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks}, author={Nils Reimers and Iryna Gurevych}, year={2019}, eprint={1908.10084}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/1908.10084}, } texttext-ranking1K<n<10K0 likes20 downloads9mo agoHugging Face07pxovela /training_setting_burnt_unet_and_text_encodertabularn<1K0 likes17 downloads3y agoHugging Face08Somnia /GPT3-Token-Encodertabularn<1K0 likes13 downloads4y agoHugging Face09GbrlOl /dataset_cross_encoder_geotechnical_report_v1.0.0 Geotechnical Reports text1K<n<10K0 likes12 downloads2y agoHugging Face10chungimungi /msmarco_hard_negatives_cross-encoderThe data was used in the paper Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval. Hard-Negatives generated by a cross-encoder using 10,000 passages from the MS-Marco dataset If this dataset was useful consider citing us :) @misc{sinha2025dontretrievegenerateprompting, title={Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval}, author={Aarush Sinha}, year={2025}, eprint={2504.21015}… See the full description on the dataset page: https://huggingface.co/datasets/chungimungi/msmarco_hard_negatives_cross-encoder.texttext-ranking10K<n<100K0 likes8 downloads9mo agoHugging Face11chungimungi /ms-marco-cross-encoder-hard-negativestext100K<n<1M0 likes8 downloads5mo agoHugging Face12Naimmm /relational_encodertext1K<n<10K0 likes6 downloads2y agoHugging Face13MLxmert0usyd1 /gsm8k-basin-encodertabularn<1K0 likes6 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.