Team Ai
20 results

embedding-models

HFforLegal /embedding-models Reference models for integration into HF for Legal 🤗 This dataset comprises a collection of models aimed at streamlining and partially automating the embedding process. Each model entry within this dataset includes essential information such as model identifiers, embedding configurations, and specific parameters, ensuring that users can seamlessly integrate these models into their workflows with minimal setup and maximum efficiency. Dataset Structure Field Type… See the full description on the dataset page: https://huggingface.co/datasets/HFforLegal/embedding-models.tabulartabular-to-textn<1K4 likes76 downloads2y agoHugging FaceLLM-OS-Models /korean-embedding-performance-v1-performance-1m Korean Embedding Performance v1 — 1M Qwen3-Embedding-8B의 한국어 retrieval data-scale 실험을 위한 정확히 1,000,000-row 연구·비상업 contrastive dataset이다. release_eligible: false, 통합 라이선스 other이며 upstream source 조건을 재허가하지 않는다. 구성 계열 Rows 비율 역할 nlpai-lab/ko-triplet-v1.0 600,254 60.03% 넓은 한국어 QA/retrieval core F2 Korean QA/instruction 287,000 28.70% webfaq, mqa, koalpaca, realQA, komagpie F2 retrieval task train-family 4,146 0.41% MIRACL, MrTidy, MLDR F2 PAWS-X… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-performance-1m.textsentence-similarity1M<n<10M0 likes73 downloads3mo agoHugging FaceMarxistLeninist /repro-how-can-embedding-models-bind-concepts-bundle# Reproduction bundle: How can embedding models bind concepts? (lFkGJ60bGq) Artifacts for the Trackio logbook space: https://huggingface.co/spaces/MarxistLeninist/repro-how-can-embedding-models-bind-concepts Paper: How can embedding models bind concepts? (ICML 2026 challenge org ICML-2026-agent-repro, orid lFkGJ60bGq, arXiv 2605.31503) Contents repro-bundle.zip / staged files: experiment scripts (bind_*.py), run logs (bind_*.log), result JSONs (results_bind_*.json), and the… See the full description on the dataset page: https://huggingface.co/datasets/MarxistLeninist/repro-how-can-embedding-models-bind-concepts-bundle.0 likes63 downloads3mo agoHugging FaceLLM-OS-Models /korean-embedding-performance-v1-sionic-retrieval-train-family-4146 Korean Sionic Retrieval Train-Family 4,146 F2LLM-v2가 공개한 Korean MIRACL, MrTidy, MLDR train-family row만 1M decontaminated curriculum에서 lossless 추출한 target-adaptation dataset이다. 공개 evaluation query는 포함하지 않으며 current-student HN7 mining 전의 source artifact다. 구성과 목적 source rows 역할 f2_miracl_ko_train 700 MIRACL Korean retrieval train-family f2_mrtidy_korean_train 1,200 MrTidy Korean train f2_mldr_ko_train 2,246 MLDR Korean long-document train-family 합계 4… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-sionic-retrieval-train-family-4146.texttext-retrieval1K<n<10K0 likes53 downloads3mo agoHugging FaceLLM-OS-Models /korean-embedding-performance-v1-pilot-50k Korean Embedding Performance v1 — Pilot 50K 주의: 이 revision은 공개 benchmark 성능 후보 학습에 사용하면 안 된다. 사후 15-task exact text-hash 감사에서 평가 query 고유 hash 4개가 확인됐다. 파이프라인·최적화 진단과 contamination ablation에만 남기며, 교체본은 ablation-200k이다. Qwen3-Embedding 계열의 한국어 retrieval 성능 실험을 위한 50,000-row 연구용 contrastive dataset이다. 각 row는 instruction-aware query, positive passage 1개, hard/easy negative passage 1–7개를 ms-swift embedding message schema로 저장한다. 사용 조건과 공개 범위 이 저장소의 통합 라이선스는 other다.… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-pilot-50k.texttext-retrieval10K<n<100K0 likes40 downloads3mo agoHugging FaceLLM-OS-Models /korean-embedding-performance-v1-ablation-200k Korean Embedding Performance v1 — Ablation 200K Qwen3-Embedding-8B의 한국어 retrieval continued fine-tuning에서 LoRA/DoRA/부분 및 full fine-tuning, loss, hard-negative 전략을 비교하기 위한 200,000-row 연구·비상업 성능 데이터다. release_eligible: false이며 통합 라이선스는 other다. upstream source별 조건을 재허가하지 않는다. 구성 계열 Rows 역할 nlpai-lab/ko-triplet-v1.0@1f5d72d 100,254 넓은 한국어 QA/retrieval core F2 Korean QA/instruction 68,000 webfaq, mqa, koalpaca, realQA, komagpie F2 retrieval task… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-ablation-200k.textsentence-similarity100K<n<1M0 likes38 downloads3mo agoHugging Face