Team Ai
Datasetpublic

CedRuiz/data_ClassificationModel

Dataset Card for "data_ClassificationModel"--- dataset_info: features: - name: brands dtype: string - name: categories dtype: string - name: code dtype: string - name: languages_tags dtype: string - name: last_modified_t dtype: int64 - name: product_name_de dtype: string - name: quantity dtype: string - name: index_level_0 dtype: int64 splits: - name: train num_bytes: 231023 num_examples: 673… See the full description on the dataset page: https://huggingface.co/datasets/CedRuiz/data_ClassificationModel.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes133downloads
check_for_duplicates_txt_data.py25 linesDownload Raw Back to root
1import os, json, hashlib2 3TEXT_DIR = "./test.txt ( Data GVK)"4unique_hashes = set()5duplicates = []6 7for fname in os.listdir(TEXT_DIR):8    if not fname.endswith(".txt"):9        continue10    try:11        with open(os.path.join(TEXT_DIR, fname), "r", encoding="utf-8") as f:12            data = json.load(f)13        text = data["responses"][0]["textAnnotations"][0]["description"].strip()14        hash_ = hashlib.md5(text.encode("utf-8")).hexdigest()15        if hash_ in unique_hashes:16            duplicates.append(fname)17        else:18            unique_hashes.add(hash_)19    except Exception:20        continue21 22print(f"Total files: {len(os.listdir(TEXT_DIR))}")23print(f"Unique OCR texts: {len(unique_hashes)}")24print(f"Duplicates found: {len(duplicates)}")25