codeShare/lora-training-data
2785
1{2 "cells": [3 {4 "cell_type": "code",5 "execution_count": null,6 "metadata": {7 "id": "_gXEa0VE88VQ"8 },9 "outputs": [],10 "source": [11 "from google.colab import drive\n",12 "drive.mount('/content/drive')"13 ]14 },15 {16 "cell_type": "markdown",17 "source": [18 "# 🚀 Authenticate credentials for Reddit PRAW into Google colab Secrets"19 ],20 "metadata": {21 "id": "gSErGKBctoAc"22 }23 },24 {25 "cell_type": "markdown",26 "source": [27 "\n",28 "\n",29 "**One-time setup** — after this, your notebook will automatically load Reddit credentials without prompts or hard-coded secrets.\n",30 "\n",31 "---\n",32 "\n",33 "### Step 1: Create Your Reddit “Script” App\n",34 "1. Go to: **[https://www.reddit.com/prefs/apps](https://www.reddit.com/prefs/apps)** (log in first)\n",35 "2. Scroll to the bottom → click **“create another app”** (or **“create app”**)\n",36 "3. Fill the form:\n",37 " - **Name**: `Colab-PRAW-Script` (or any name you like)\n",38 " - **App type**: Select **script** ← very important\n",39 " - **Description**: (optional)\n",40 " - **Redirect URI**: `http://localhost:8080`\n",41 "4. Click **Create app**\n",42 "\n",43 "You will now see:\n",44 "- **personal use script** → 14-character string → **REDDIT_CLIENT_ID**\n",45 "- **secret** → long string → **REDDIT_CLIENT_SECRET**\n",46 "\n",47 "Copy both values.\n",48 "\n",49 "---\n",50 "\n",51 "### Step 2: Choose Your User-Agent\n",52 "Use this format (replace `yourusername` with your actual Reddit username):"53 ],54 "metadata": {55 "id": "rst5O8t0tW7F"56 }57 },58 {59 "cell_type": "markdown",60 "source": [61 ""62 ],63 "metadata": {64 "id": "rQPyne1Hr_CM"65 }66 },67 {68 "cell_type": "markdown",69 "source": [70 "This will be your **REDDIT_USER_AGENT**.\n",71 "\n",72 "---\n",73 "\n",74 "### Step 3: Add Secrets in Google Colab\n",75 "1. Open your Colab notebook\n",76 "2. In the left sidebar click the **🔑 Secrets** tab\n",77 "3. Click **+ Add new secret** for each of these:\n",78 "\n",79 "| Secret Name | What to paste |\n",80 "|------------------------|----------------------------------------------------|\n",81 "| `REDDIT_CLIENT_ID` | 14-character string from Reddit |\n",82 "| `REDDIT_CLIENT_SECRET` | The secret string from Reddit |\n",83 "| `REDDIT_USER_AGENT` | `Colab-PRAW:v1.0 (by u/yourusername)` |\n",84 "| `REDDIT_USERNAME` | Your Reddit username (no `u/`) |\n",85 "| `REDDIT_PASSWORD` | Your Reddit account password |\n",86 "\n",87 "After each one, click **Add** and **Grant access** when prompted.\n",88 "\n",89 "---"90 ],91 "metadata": {92 "id": "d4ikd1h7tbr5"93 }94 },95 {96 "cell_type": "code",97 "source": [98 "#@markdown **Reddit PRAW Authentication with Colab Secrets**\n",99 "\n",100 "# Install PRAW (run once)\n",101 "!pip install praw -q\n",102 "\n",103 "import praw\n",104 "from google.colab import userdata\n",105 "import getpass\n",106 "\n",107 "# === AUTHENTICATION (uses secrets first, falls back to prompt if missing) ===\n",108 "def get_secret_or_prompt(key, prompt_text):\n",109 " try:\n",110 " return userdata.get(key)\n",111 " except (KeyError, Exception):\n",112 " return getpass.getpass(prompt_text)\n",113 "\n",114 "reddit = praw.Reddit(\n",115 " client_id=get_secret_or_prompt(\"REDDIT_CLIENT_ID\", \"Enter your REDDIT_CLIENT_ID: \"),\n",116 " client_secret=get_secret_or_prompt(\"REDDIT_CLIENT_SECRET\", \"Enter your REDDIT_CLIENT_SECRET: \"),\n",117 " user_agent=get_secret_or_prompt(\"REDDIT_USER_AGENT\", \"Enter your REDDIT_USER_AGENT: \"),\n",118 " username=get_secret_or_prompt(\"REDDIT_USERNAME\", \"Enter your REDDIT_USERNAME: \"),\n",119 " password=get_secret_or_prompt(\"REDDIT_PASSWORD\", \"Enter your REDDIT_PASSWORD: \"),\n",120 ")\n",121 "\n",122 "# Quick verification\n",123 "print(\"✅ Successfully logged in as:\", reddit.user.me())\n",124 "print(\"🔒 Read-only mode:\", reddit.read_only)"125 ],126 "metadata": {127 "cellView": "form",128 "id": "WLopd-NFtHwY"129 },130 "execution_count": null,131 "outputs": []132 },133 {134 "cell_type": "markdown",135 "metadata": {136 "id": "V5i2-xwTGG4K"137 },138 "source": [139 "#🚀 Fetch content from Reddit"140 ]141 },142 {143 "cell_type": "markdown",144 "source": [145 "# Image crop left rigt"146 ],147 "metadata": {148 "id": "WX17R9Yc6U1A"149 }150 },151 {152 "cell_type": "code",153 "execution_count": null,154 "metadata": {155 "id": "-ewlMjrTga21",156 "cellView": "form"157 },158 "outputs": [],159 "source": [160 "# === INTEGRATED REDDIT DOWNLOADER + MEDIA EXTRACTOR (Drive + Checkboxes + Sliders Edition) ===\n",161 "# Replace BOTH of your previous cells with this single block and run it.\n",162 "# All results (titles, links, images, frames, zips) are saved directly to your Google Drive.\n",163 "# Images/frames are numbered sequentially (1.jpeg, 2.jpeg, ...) with optional paired .txt files.\n",164 "\n",165 "# @markdown ** Subreddit & Fetch Settings**\n",166 "subreddit_name = \"celebnsfw\" # @param {type:\"string\"}\n",167 "sort_method = \"new\" # @param [\"hot\", \"new\", \"top\"] {type:\"string\"}\n",168 "num_posts_to_pull = 100 # @param {type:\"slider\", min:0, max:2000, step:100}\n",169 "offset_index = 0 # @param {type:\"slider\", min:0, max:2000, step:100}\n",170 "\n",171 "# @markdown **✅ Functionality Checkboxes (select any combination)**\n",172 "thumbnail_low_res = True # @param {type:\"boolean\"}\n",173 "#1) Thumbnail download at low res\n",174 "proper_image_download = False # @param {type:\"boolean\"}\n",175 "#2) Proper image download including sub images in galleries\n",176 "gif_frame_extraction = False # @param {type:\"boolean\"}\n",177 "#3) GIF + RedGifs video → keyframe extraction (original files are NEVER saved)\n",178 "combined_titles_txt = True # @param {type:\"boolean\"}\n",179 "#4) Combined txt file of titles\n",180 "pair_txt_with_media = False # @param {type:\"boolean\"}\n",181 "#5) Adding txt files to saved images in zips as enumerated pairs (1.txt, 2.txt, ...)\n",182 "debug_mode = False # @param {type:\"boolean\"}\n",183 "#6) Enable detailed debug printouts during media processing\n",184 "\n",185 "import os\n",186 "import shutil\n",187 "import glob\n",188 "import requests\n",189 "import subprocess\n",190 "from google.colab import files, drive\n",191 "\n",192 "if gif_frame_extraction:\n",193 " # Install yt-dlp if it's not already installed (moved outside conditional block)\n",194 " !pip install -q yt-dlp\n",195 " import yt_dlp\n",196 "\n",197 "\n",198 "# ========================== DRIVE SETUP ==========================\n",199 "print(\" Mounting Google Drive...\")\n",200 "drive.mount('/content/drive', force_remount=False)\n",201 "\n",202 "drive_base_dir = f\"/content/drive/MyDrive/{subreddit_name}_reddit_downloads\"\n",203 "os.makedirs(drive_base_dir, exist_ok=True)\n",204 "\n",205 "# Local temp media storage (Not on Drive)\n",206 "local_media_dir = f\"/content/temp_output/{subreddit_name}_media\"\n",207 "if os.path.exists(local_media_dir): shutil.rmtree(local_media_dir)\n",208 "os.makedirs(local_media_dir, exist_ok=True)\n",209 "\n",210 "local_output_dir = \"/content/output\"\n",211 "os.makedirs(local_output_dir, exist_ok=True)\n",212 "\n",213 "print(f\" Individual images stay local; only Zips/Txt go to: {drive_base_dir}\")\n",214 "\n",215 "# ========================== FETCH POSTS ==========================\n",216 "print(f\" Fetching up to {num_posts_to_pull} posts from r/{subreddit_name} ({sort_method} sorting, offset {offset_index})...\")\n",217 "\n",218 "sub = reddit.subreddit(subreddit_name)\n",219 "\n",220 "if sort_method == \"hot\":\n",221 " iterator = sub.hot(limit=offset_index + num_posts_to_pull + 50)\n",222 "elif sort_method == \"new\":\n",223 " iterator = sub.new(limit=offset_index + num_posts_to_pull + 50)\n",224 "elif sort_method == \"top\":\n",225 " iterator = sub.top(limit=offset_index + num_posts_to_pull + 50)\n",226 "else:\n",227 " iterator = sub.hot(limit=offset_index + num_posts_to_pull + 50)\n",228 "\n",229 "all_posts = list(iterator)\n",230 "posts = all_posts[offset_index : offset_index + num_posts_to_pull]\n",231 "\n",232 "print(f\"✅ Successfully fetched {len(posts)} posts.\")\n",233 "\n",234 "# === NEW: Filename suffix with actual fetched count + offset ===\n",235 "fetch_suffix = f\"_{len(posts)}posts_offset{offset_index}\"\n",236 "\n",237 "titles = []\n",238 "external_links = []\n",239 "\n",240 "for submission in posts:\n",241 " cleaned_title = (submission.title\n",242 " .replace('^', '').replace('{', '').replace('}', '').replace('|', '')\n",243 " .replace('[','').replace(']','').replace('\"',''))\n",244 " titles.append(cleaned_title)\n",245 "\n",246 " url = getattr(submission, 'url', None)\n",247 " if (url and url.startswith(('http://', 'https://')) and\n",248 " not any(domain in url for domain in ['reddit.com', 'redd.it'])):\n",249 " external_links.append(url)\n",250 "\n",251 "# ========================== OPTION 4: Combined Titles TXT ==========================\n",252 "if combined_titles_txt:\n",253 " combined_content = '{' + '|'.join(titles) + '}'\n",254 " titles_file_drive = f\"{drive_base_dir}/{subreddit_name}_titles{fetch_suffix}.txt\"\n",255 " with open(titles_file_drive, \"w\", encoding=\"utf-8\") as f:\n",256 " f.write(combined_content)\n",257 " print(f\" Combined titles saved → {titles_file_drive}\")\n",258 "\n",259 "# ========================== External Links ==========================\n",260 "if external_links:\n",261 " links_file = f\"{drive_base_dir}/{subreddit_name}_links{fetch_suffix}.txt\"\n",262 " with open(links_file, \"w\", encoding=\"utf-8\") as f:\n",263 " for link in external_links: f.write(link + \"\\n\")\n",264 " print(f\" External links saved → {links_file}\")\n",265 " print(f\" → {len(external_links)} links (mostly RedGifs)\")\n",266 "\n",267 "# ========================== MEDIA PROCESSING ==========================\n",268 "any_media = thumbnail_low_res or proper_image_download or gif_frame_extraction\n",269 "\n",270 "if any_media:\n",271 "\n",272 " temp_dir = \"/content/temp_download\"\n",273 " os.makedirs(temp_dir, exist_ok=True)\n",274 "\n",275 " global_counter = 1\n",276 "\n",277 " def save_media_and_txt(source, is_url=False, title=\"\"):\n",278 " global global_counter\n",279 " if is_url:\n",280 " clean_url = source.split('?')[0]\n",281 " ext = os.path.splitext(clean_url)[1].lower()\n",282 " if ext not in ['.jpg', '.jpeg', '.png']: ext = '.jpg'\n",283 " temp_file = f\"{temp_dir}/dl_{global_counter}{ext}\"\n",284 " try:\n",285 " r = requests.get(source, stream=True, timeout=60)\n",286 " r.raise_for_status()\n",287 " with open(temp_file, 'wb') as f:\n",288 " for chunk in r.iter_content(8192): f.write(chunk)\n",289 " except Exception:\n",290 " return\n",291 " else:\n",292 " temp_file = source\n",293 " ext = os.path.splitext(temp_file)[1].lower() or '.jpeg'\n",294 "\n",295 " local_path = f\"{local_media_dir}/{global_counter}{ext}\"\n",296 " shutil.copy2(temp_file, local_path)\n",297 "\n",298 " if pair_txt_with_media:\n",299 " with open(f\"{local_media_dir}/{global_counter}.txt\", \"w\", encoding=\"utf-8\") as f:\n",300 " f.write(title)\n",301 "\n",302 " global_counter += 1\n",303 " print(f\" ✅ Saved locally #{global_counter-1} \", end=\"\\r\")\n",304 " if is_url and os.path.exists(temp_file): os.remove(temp_file)\n",305 "\n",306 " for idx, submission in enumerate(posts, 1):\n",307 " cleaned_title = titles[idx-1]\n",308 "\n",309 " if thumbnail_low_res:\n",310 " thumb_url = getattr(submission, 'thumbnail', None)\n",311 " if thumb_url and thumb_url.startswith('http'):\n",312 " save_media_and_txt(thumb_url, is_url=True, title=cleaned_title)\n",313 "\n",314 " if proper_image_download:\n",315 " url = getattr(submission, 'url', None)\n",316 " if url and any(url.lower().endswith(e) for e in ['.jpg', '.jpeg', '.png']):\n",317 " save_media_and_txt(url, is_url=True, title=cleaned_title)\n",318 "\n",319 " # === GIF + RedGifs keyframe extraction (debug toggleable) ===\n",320 " if gif_frame_extraction:\n",321 " url = getattr(submission, 'url', None)\n",322 " if url and url.startswith(('http://', 'https://')):\n",323 " if debug_mode:\n",324 " print(f\"\u001f DEBUG [{idx:03d}]: Checking URL → {url[:90]}...\")\n",325 "\n",326 " if url.lower().endswith('.gif') or 'redgifs.com' in url.lower():\n",327 " if debug_mode:\n",328 " print(f\"✅ DEBUG: Eligible for keyframe extraction → {url[:90]}...\")\n",329 " temp_media = None\n",330 " try:\n",331 " # 1. Download\n",332 " if url.lower().endswith('.gif'):\n",333 " temp_media = f\"{temp_dir}/dl_gif_{idx}.gif\"\n",334 " r = requests.get(url, stream=True, timeout=60)\n",335 " r.raise_for_status()\n",336 " with open(temp_media, 'wb') as f:\n",337 " for chunk in r.iter_content(8192):\n",338 " f.write(chunk)\n",339 " if debug_mode:\n",340 " print(f\" ↓ Downloaded GIF ({os.path.getsize(temp_media)/1024/1024:.1f} MB)\")\n",341 " else:\n",342 " # RedGifs video\n",343 " temp_media = f\"{temp_dir}/dl_video_{idx}.mp4\"\n",344 " ydl_opts = {\n",345 " 'outtmpl': temp_media,\n",346 " 'quiet': True,\n",347 " 'no_warnings': True,\n",348 " 'format': 'bestvideo+bestaudio/best',\n",349 " 'merge_output_format': 'mp4',\n",350 " }\n",351 " with yt_dlp.YoutubeDL(ydl_opts) as ydl:\n",352 " ydl.download([url])\n",353 " if debug_mode:\n",354 " print(f\" ↓ Downloaded RedGifs video ({os.path.getsize(temp_media)/1024/1024:.1f} MB)\")\n",355 "\n",356 " # 2. Keyframe extraction\n",357 " output_pattern = f\"{temp_dir}/kf_{idx}_%04d.jpg\"\n",358 " cmd = [\n",359 " 'ffmpeg', '-y', '-i', temp_media,\n",360 " '-vf', \"select='gt(scene,0.15)',setpts=N/(FRAME_RATE*TB)\",\n",361 " '-vsync', 'vfr',\n",362 " '-q:v', '5',\n",363 " output_pattern\n",364 " ]\n",365 " subprocess.run(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE, check=False)\n",366 " if debug_mode:\n",367 " print(f\" 🏆 ffmpeg keyframe extraction finished (scene threshold 0.15)\")\n",368 "\n",369 " # 3. Save keyframes\n",370 " extracted_frames = sorted(glob.glob(f\"{temp_dir}/kf_{idx}_*.jpg\"))\n",371 " if debug_mode:\n",372 " print(f\" 📎 Extracted {len(extracted_frames)} keyframes → saving to ZIP\")\n",373 "\n",374 " for frame_path in extracted_frames:\n",375 " save_media_and_txt(frame_path, is_url=False, title=cleaned_title)\n",376 " os.remove(frame_path)\n",377 "\n",378 " # 4. Cleanup original\n",379 " if temp_media and os.path.exists(temp_media):\n",380 " os.remove(temp_media)\n",381 " if debug_mode:\n",382 " print(f\" 🗑️ Cleaned up original media file\")\n",383 "\n",384 " except Exception as e:\n",385 " if debug_mode:\n",386 " print(f\" ⚠️ ERROR processing {url[:80]}... → {e}\")\n",387 " if temp_media and os.path.exists(temp_media):\n",388 " os.remove(temp_media)\n",389 " continue\n",390 "\n",391 " print(f\"\\n\\n✅ MEDIA CACHED! Creating ZIP on Drive...\")\n",392 " zip_name = f\"{subreddit_name}_media{fetch_suffix}\"\n",393 " shutil.make_archive(f\"{drive_base_dir}/{zip_name}\", 'zip', local_media_dir)\n",394 " print(f\" ZIP created on Drive → {drive_base_dir}/{zip_name}.zip\")\n",395 "\n",396 "print(\"\\n✨ ALL DONE!\")"397 ]398 },399 {400 "cell_type": "markdown",401 "metadata": {402 "id": "3b94280b"403 },404 "source": [405 "### 🖼️ Image Cropper: 1024x1024 Left and Right Aligned"406 ]407 },408 {409 "cell_type": "code",410 "metadata": {411 "id": "d3ca4317"412 },413 "source": [414 "from google.colab import drive\n",415 "from PIL import Image\n",416 "import os\n",417 "import shutil\n",418 "import glob\n",419 "import zipfile\n",420 "\n",421 "# @markdown **Configuration**\n",422 "input_image_directory = \"/content\" # @param {type:\"string\"}\n",423 "output_drive_folder = \"cropped_1024x1024_images\" # @param {type:\"string\"}\n",424 "target_size = 1024 # @param {type:\"integer\"}\n",425 "resize_smaller_images = True # @param {type:\"boolean\"}\n",426 "\n",427 "# Mount Google Drive\n",428 "print(\"🔄 Mounting Google Drive...\")\n",429 "drive.mount('/content/drive', force_remount=True)\n",430 "\n",431 "# Define output paths\n",432 "drive_base_path = '/content/drive/MyDrive'\n",433 "drive_output_folder_path = os.path.join(drive_base_path, output_drive_folder)\n",434 "\n",435 "# Create a temporary local directory for cropped images before zipping\n",436 "local_temp_crop_dir = '/content/temp_cropped_images'\n",437 "os.makedirs(local_temp_crop_dir, exist_ok=True)\n",438 "\n",439 "print(f\"\\n🔍 Scanning for images in {input_image_directory}...\")\n",440 "image_files = []\n",441 "for ext in ('*.png', '*.jpg', '*.jpeg', '*.webp'):\n",442 " image_files.extend(glob.glob(os.path.join(input_image_directory, ext)))\n",443 "\n",444 "if not image_files:\n",445 " print(f\"❌ No image files found in {input_image_directory}. Please ensure images are present or check the directory path.\")\n",446 "else:\n",447 " print(f\"✅ Found {len(image_files)} image(s) to process.\")\n",448 " processed_count = 0\n",449 " skipped_count = 0\n",450 "\n",451 " for file_path in image_files:\n",452 " filename = os.path.basename(file_path)\n",453 " print(f\"\\nProcessing {filename}...\")\n",454 " try:\n",455 " img = Image.open(file_path)\n",456 "\n",457 " if img.mode != 'RGB':\n",458 " img = img.convert('RGB')\n",459 "\n",460 " width, height = img.size\n",461 "\n",462 " if width < target_size or height < target_size:\n",463 " if resize_smaller_images:\n",464 " print(f\"⚠️ Image dimensions ({width}x{height}) are smaller than target {target_size}x{target_size}. Resizing...\")\n",465 " if width < height:\n",466 " new_width = target_size\n",467 " new_height = int(height * (target_size / width))\n",468 " else:\n",469 " new_height = target_size\n",470 " new_width = int(width * (target_size / height))\n",471 " img = img.resize((new_width, new_height), Image.LANCZOS)\n",472 " width, height = img.size\n",473 " print(f\" Resized to {width}x{height}.\")\n",474 " else:\n",475 " print(f\"⚠️ Skipping {filename}: Image dimensions ({width}x{height}) are smaller than target {target_size}x{target_size}.\")\n",476 " skipped_count += 1\n",477 " continue\n",478 "\n",479 " # --- Left-aligned crop ---\n",480 " left_crop_box = (0, 0, target_size, target_size)\n",481 " img_left_cropped = img.crop(left_crop_box)\n",482 " output_name_left = f\"{os.path.splitext(filename)[0]}_left_{target_size}x{target_size}.jpg\"\n",483 " img_left_cropped.save(os.path.join(local_temp_crop_dir, output_name_left))\n",484 " print(f\" ✅ Saved left-aligned crop to temp: {output_name_left}\")\n",485 "\n",486 " # --- Right-aligned crop ---\n",487 " right_x_start = width - target_size\n",488 " right_crop_box = (right_x_start, 0, width, target_size)\n",489 " img_right_cropped = img.crop(right_crop_box)\n",490 " output_name_right = f\"{os.path.splitext(filename)[0]}_right_{target_size}x{target_size}.jpg\"\n",491 " img_right_cropped.save(os.path.join(local_temp_crop_dir, output_name_right))\n",492 " print(f\" ✅ Saved right-aligned crop to temp: {output_name_right}\")\n",493 "\n",494 " processed_count += 1\n",495 "\n",496 " except Exception as e:\n",497 " print(f\"❌ Error processing {filename}: {e}\")\n",498 " skipped_count += 1\n",499 "\n",500 " print(f\"\\n✨ Processing complete! {processed_count} images processed, {skipped_count} skipped.\")\n",501 "\n",502 " # Create ZIP file\n",503 " if processed_count > 0:\n",504 " os.makedirs(drive_output_folder_path, exist_ok=True)\n",505 " zip_filename = f\"{output_drive_folder}.zip\"\n",506 " zip_path_on_drive = os.path.join(drive_output_folder_path, zip_filename)\n",507 "\n",508 " print(f\"\\n📦 Creating ZIP file '{zip_filename}' in Google Drive...\")\n",509 " shutil.make_archive(os.path.join(drive_output_folder_path, output_drive_folder), 'zip', local_temp_crop_dir)\n",510 " print(f\"✅ All cropped images zipped and saved to: {zip_path_on_drive}\")\n",511 " else:\n",512 " print(\"No images were processed to be zipped.\")\n",513 "\n",514 " # Clean up local temporary directory\n",515 " shutil.rmtree(local_temp_crop_dir)\n",516 " print(\"🗑️ Cleaned up local temporary cropped images directory.\")\n"517 ],518 "execution_count": null,519 "outputs": []520 },521 {522 "cell_type": "markdown",523 "metadata": {524 "id": "qisU9VeIyzX2"525 },526 "source": [527 "# 🎮 itch.io (Indie Game Website) fetch"528 ]529 },530 {531 "cell_type": "code",532 "execution_count": null,533 "metadata": {534 "id": "HDCueX6d3K00",535 "cellView": "form"536 },537 "outputs": [],538 "source": [539 "#@markdown # 🐙 itch.io Image-Text Dataset Creator\n",540 "#@markdown **Fetch enumerated thumbnails + matching TXT files → ZIP → Google Drive**\n",541 "\n",542 "#@markdown ---\n",543 "#@markdown ### 📋 Choose your settings below then **Run this cell**\n",544 "\n",545 "num_games = 1000 #@param {type:\"slider\", min:10, max:5000, step:10, description:\"How many games to fetch\"}\n",546 "\n",547 "sort_by = \"Most Recent\" #@param [\"Popular (default)\", \"New & Popular\", \"Top Sellers\", \"Top Rated\", \"Most Recent\"]\n",548 "\n",549 "subcategory = \"None (All Games)\" #@param [\"None (All Games)\", \"Genre: Action\", \"Genre: Adventure\", \"Genre: Arcade\", \"Genre: Card Game\", \"Genre: Educational\", \"Genre: Fighting\", \"Genre: Platformer\", \"Genre: Puzzle\", \"Genre: RPG\", \"Genre: Shooter\", \"Genre: Simulation\", \"Genre: Strategy\", \"Genre: Visual Novel\", \"Platform: Web\", \"Platform: Windows\", \"Platform: macOS\", \"Platform: Linux\", \"Platform: Android\", \"Tag: 2D\", \"Tag: Pixel Art\", \"Tag: Horror\", \"Tag: Multiplayer\", \"Tag: Roguelike\", \"Tag: Retro\"]\n",550 "\n",551 "#@markdown ---\n",552 "\n",553 "#@markdown **After changing the values above, just click the ▶️ Run button on this cell.**\n",554 "\n",555 "# ================================================\n",556 "# ✅ READY-TO-RUN COLAB CELL (single cell version)\n",557 "# ================================================\n",558 "\n",559 "import requests\n",560 "from bs4 import BeautifulSoup\n",561 "import os\n",562 "import re\n",563 "import json\n",564 "import time\n",565 "import shutil\n",566 "import datetime\n",567 "from urllib.parse import urljoin\n",568 "from IPython.display import display, HTML, Image\n",569 "from google.colab import drive\n",570 "\n",571 "print(\"✅ Starting itch.io scraper with your chosen settings...\")\n",572 "\n",573 "# ====================== MAP USER CHOICES TO URL SLUGS ======================\n",574 "sort_map = {\n",575 " \"Popular (default)\": \"\",\n",576 " \"New & Popular\": \"new-and-popular\",\n",577 " \"Top Sellers\": \"top-sellers\",\n",578 " \"Top Rated\": \"top-rated\",\n",579 " \"Most Recent\": \"newest\"\n",580 "}\n",581 "\n",582 "filter_map = {\n",583 " \"None (All Games)\": \"\",\n",584 " \"Genre: Action\": \"genre-action\",\n",585 " \"Genre: Adventure\": \"genre-adventure\",\n",586 " \"Genre: Arcade\": \"genre-arcade\",\n",587 " \"Genre: Card Game\": \"genre-card-game\",\n",588 " \"Genre: Educational\": \"genre-educational\",\n",589 " \"Genre: Fighting\": \"genre-fighting\",\n",590 " \"Genre: Platformer\": \"genre-platformer\",\n",591 " \"Genre: Puzzle\": \"genre-puzzle\",\n",592 " \"Genre: RPG\": \"genre-rpg\",\n",593 " \"Genre: Shooter\": \"genre-shooter\",\n",594 " \"Genre: Simulation\": \"genre-simulation\",\n",595 " \"Genre: Strategy\": \"genre-strategy\",\n",596 " \"Genre: Visual Novel\": \"genre-visual-novel\",\n",597 " \"Platform: Web\": \"platform-web\",\n",598 " \"Platform: Windows\": \"platform-windows\",\n",599 " \"Platform: macOS\": \"platform-macos\",\n",600 " \"Platform: Linux\": \"platform-linux\",\n",601 " \"Platform: Android\": \"platform-android\",\n",602 " \"Tag: 2D\": \"tag-2d\",\n",603 " \"Tag: Pixel Art\": \"tag-pixel-art\",\n",604 " \"Tag: Horror\": \"tag-horror\",\n",605 " \"Tag: Multiplayer\": \"tag-multiplayer\",\n",606 " \"Tag: Roguelike\": \"tag-roguelike\",\n",607 " \"Tag: Retro\": \"tag-retro\"\n",608 "}\n",609 "\n",610 "sort_slug = sort_map[sort_by]\n",611 "filter_slug = filter_map[subcategory]\n",612 "\n",613 "# ====================== CONFIG ======================\n",614 "MAX_PAGES = 20\n",615 "DELAY_BETWEEN_PAGES = 1.5\n",616 "HEADERS = {\n",617 " \"User-Agent\": \"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 \"\n",618 " \"(KHTML, like Gecko) Chrome/134.0 Safari/537.36\"\n",619 "}\n",620 "\n",621 "# Build the base URL according to user selections\n",622 "base_url = \"https://itch.io/games\"\n",623 "if sort_slug:\n",624 " base_url += f\"/{sort_slug}\"\n",625 "if filter_slug:\n",626 " base_url += f\"/{filter_slug}\"\n",627 "\n",628 "print(f\"🌐 Target: {base_url} | Fetching up to {num_games} games\")\n",629 "\n",630 "# Mount Google Drive\n",631 "print(\"🔄 Mounting Google Drive...\")\n",632 "drive.mount('/content/drive', force_remount=True)\n",633 "\n",634 "# Create working folder\n",635 "dataset_dir = \"/content/itch_dataset\"\n",636 "os.makedirs(dataset_dir, exist_ok=True)\n",637 "\n",638 "# ====================== SCRAPE MULTIPLE PAGES ======================\n",639 "pairs = []\n",640 "page = 1\n",641 "\n",642 "while len(pairs) < num_games and page <= MAX_PAGES:\n",643 " url = f\"{base_url}?page={page}\" if page > 1 else base_url\n",644 " print(f\"📄 Scraping page {page} → {url}\")\n",645 "\n",646 " try:\n",647 " response = requests.get(url, headers=HEADERS, timeout=20)\n",648 " if response.status_code != 200:\n",649 " print(f\"❌ Page {page} failed (HTTP {response.status_code})\")\n",650 " break\n",651 "\n",652 " soup = BeautifulSoup(response.text, \"html.parser\")\n",653 " game_cells = soup.find_all(\"div\", class_=\"game_cell\")\n",654 "\n",655 " if not game_cells:\n",656 " print(\"✅ No more games on this page.\")\n",657 " break\n",658 "\n",659 " added = 0\n",660 " for cell in game_cells:\n",661 " if len(pairs) >= num_games:\n",662 " break\n",663 "\n",664 " # Clean title (no price spam)\n",665 " title_tag = cell.find(\"div\", class_=\"game_title\")\n",666 " if not title_tag:\n",667 " continue\n",668 " title_link = title_tag.find(\"a\", class_=\"title\")\n",669 " title = title_link.get_text(strip=True) if title_link else title_tag.get_text(strip=True).split(\"$\")[0].strip()\n",670 " if not title:\n",671 " continue\n",672 "\n",673 " # Extract thumbnail (robust for current itch.io layout)\n",674 " img_url = None\n",675 " img_tag = cell.find(\"img\")\n",676 " if img_tag:\n",677 " for attr in [\"data-lazy-src\", \"data-lazy_src\", \"data-src\", \"src\"]:\n",678 " img_url = img_tag.get(attr)\n",679 " if img_url:\n",680 " break\n",681 " if not img_url and img_tag.get(\"srcset\"):\n",682 " img_url = img_tag.get(\"srcset\").split(\",\")[0].strip().split(\" \")[0]\n",683 "\n",684 " # Fallback: background-image\n",685 " if not img_url:\n",686 " for el in cell.find_all(lambda t: t.has_attr(\"style\") and \"background-image\" in t.get(\"style\", \"\").lower()):\n",687 " style = el.get(\"style\", \"\")\n",688 " match = re.search(r'background-image\\s*:\\s*url\\([\\'\\\"]?([^\\'\\\"]+)[\\'\\\"]?\\)', style, re.IGNORECASE)\n",689 " if match:\n",690 " img_url = match.group(1)\n",691 " break\n",692 "\n",693 " if img_url:\n",694 " if img_url.startswith(\"//\"):\n",695 " img_url = \"https:\" + img_url\n",696 " elif not img_url.startswith((\"http://\", \"https://\")):\n",697 " img_url = urljoin(\"https://itch.io\", img_url)\n",698 "\n",699 " pairs.append((title, img_url))\n",700 " added += 1\n",701 "\n",702 " print(f\" → Added {added} games (total so far: {len(pairs)})\")\n",703 "\n",704 " except Exception as e:\n",705 " print(f\"❌ Error on page {page}: {e}\")\n",706 " break\n",707 "\n",708 " page += 1\n",709 " time.sleep(DELAY_BETWEEN_PAGES)\n",710 "\n",711 "if not pairs:\n",712 " print(\"❌ No games found with current filters. Try different settings.\")\n",713 "else:\n",714 " print(f\"\\n✅ Collected {len(pairs)} games. Downloading images and creating TXT files...\")\n",715 "\n",716 " # ====================== DOWNLOAD ENUMERATED FILES ======================\n",717 " downloaded = 0\n",718 " for idx, (title, img_url) in enumerate(pairs, start=1):\n",719 " num_str = f\"{idx:04d}\"\n",720 " img_path = f\"{dataset_dir}/{num_str}.jpg\"\n",721 " txt_path = f\"{dataset_dir}/{num_str}.txt\"\n",722 "\n",723 " try:\n",724 " img_response = requests.get(img_url, headers=HEADERS, timeout=15)\n",725 " if img_response.status_code == 200:\n",726 " with open(img_path, \"wb\") as f:\n",727 " f.write(img_response.content)\n",728 " with open(txt_path, \"w\", encoding=\"utf-8\") as f:\n",729 " f.write(title)\n",730 " downloaded += 1\n",731 " if downloaded % 10 == 0 or downloaded == len(pairs):\n",732 " print(f\" ✅ Saved {downloaded:04d}.jpg + {num_str}.txt\")\n",733 " else:\n",734 " print(f\"⚠️ Failed to download image {num_str}\")\n",735 " except Exception as e:\n",736 " print(f\"❌ Error downloading {num_str}: {e}\")\n",737 "\n",738 " # ====================== CREATE ZIP & SAVE TO DRIVE ======================\n",739 " timestamp = datetime.datetime.now().strftime(\"%Y%m%d_%H%M%S\")\n",740 " zip_name = f\"itch_games_{timestamp}\"\n",741 " zip_path_local = f\"/content/{zip_name}.zip\"\n",742 "\n",743 " print(f\"\\n🗜️ Creating ZIP file with {downloaded} image+txt pairs...\")\n",744 " shutil.make_archive(f\"/content/{zip_name}\", 'zip', dataset_dir)\n",745 "\n",746 " # Copy to Google Drive\n",747 " drive_folder = \"/content/drive/MyDrive/itch_datasets\"\n",748 " os.makedirs(drive_folder, exist_ok=True)\n",749 " drive_zip_path = f\"{drive_folder}/{zip_name}.zip\"\n",750 " shutil.copy(zip_path_local, drive_zip_path)\n",751 "\n",752 " print(\"\\n\" + \"=\"*80)\n",753 " print(\"🎉 SUCCESS! Your dataset is ready\")\n",754 " print(f\"📦 ZIP file: {zip_name}.zip\")\n",755 " print(f\"📤 Saved to Google Drive → {drive_zip_path}\")\n",756 " print(f\" • Files inside: 0001.jpg + 0001.txt, 0002.jpg + 0002.txt, ...\")\n",757 " print(f\" • Total pairs: {downloaded}\")\n",758 " print(\"=\"*80)\n",759 "\n",760 " # ====================== PREVIEW FIRST 3 PAIRS ======================\n",761 " print(\"\\n📸 Preview of first 3 pairs (click images to enlarge):\")\n",762 " for i in range(min(3, len(pairs))):\n",763 " num_str = f\"{i+1:04d}\"\n",764 " img_file = f\"{dataset_dir}/{num_str}.jpg\"\n",765 " display(HTML(f\"<h4>{num_str}. {pairs[i][0]}</h4>\"))\n",766 " display(Image(filename=img_file, width=400))\n",767 " print(\"─\" * 70)\n",768 "\n",769 " print(f\"\\n✅ All files are also in: {dataset_dir} (you can download the folder manually if needed)\")"770 ]771 },772 {773 "cell_type": "markdown",774 "source": [775 "# 📚 VNDB (Visual Novel Database) fetch"776 ],777 "metadata": {778 "id": "0VlyEacYp7Pn"779 }780 },781 {782 "cell_type": "code",783 "execution_count": null,784 "metadata": {785 "id": "70mWljq0EJCT",786 "cellView": "form"787 },788 "outputs": [],789 "source": [790 "#@markdown # 🐙 VNDB Image-Text fetch\n",791 "#@markdown **✅ Added offset slider + metadata in ZIP**\n",792 "#@markdown Now you can fetch any batch (e.g. first 1000 → offset 0, next 1000 → offset 1000, etc.)\n",793 "\n",794 "#@markdown ---\n",795 "#@markdown ### 📋 Choose your settings below then **Run this cell**\n",796 "\n",797 "num_vns = 1000 #@param {type:\"slider\", min:10, max:5000, step:10, description:\"How many visual novels to fetch in this batch\"}\n",798 "\n",799 "offset = 0 #@param {type:\"slider\", min:0, max:20000, step:100}\n",800 "#description:\"Offset: skip this many VNs before starting (0 = first batch, 1000 = second batch, etc.)\"}\n",801 "\n",802 "sort_by = \"Most Recent (released desc)\" #@param [\"Most Recent (released desc)\", \"Highest Rated\", \"Most Voted\"]\n",803 "\n",804 "#@markdown **Tag ID** (from your link: https://vndb.org/g3560)\n",805 "tag_id = \"g3560\" #@param {type:\"string\", description:\"VNDB tag ID (e.g. g3560 = 3D Graphics)\"}\n",806 "\n",807 "#@markdown ---\n",808 "\n",809 "#@markdown **After changing the values above, just click the ▶️ Run button on this cell.**\n",810 "\n",811 "# ================================================\n",812 "# ✅ FULLY DEBUGGED + OFFSET + METADATA READY-TO-RUN COLAB CELL\n",813 "# ================================================\n",814 "\n",815 "import requests\n",816 "import os\n",817 "import json\n",818 "import time\n",819 "import shutil\n",820 "import datetime\n",821 "from IPython.display import display, HTML, Image\n",822 "from google.colab import drive\n",823 "\n",824 "print(\"✅ Starting VNDB scraper with OFFSET support...\")\n",825 "\n",826 "# ====================== CONFIG ======================\n",827 "MAX_RESULTS_PER_PAGE = 100\n",828 "DELAY_BETWEEN_PAGES = 0.5\n",829 "HEADERS = {\n",830 " \"User-Agent\": \"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 \"\n",831 " \"(KHTML, like Gecko) Chrome/134.0 Safari/537.36\",\n",832 " \"Content-Type\": \"application/json\"\n",833 "}\n",834 "\n",835 "API_URL = \"https://api.vndb.org/kana/vn\"\n",836 "\n",837 "# Sort mapping\n",838 "sort_map = {\n",839 " \"Most Recent (released desc)\": {\"sort\": \"released\", \"reverse\": True},\n",840 " \"Highest Rated\": {\"sort\": \"rating\", \"reverse\": True},\n",841 " \"Most Voted\": {\"sort\": \"votecount\", \"reverse\": True}\n",842 "}\n",843 "\n",844 "selected_sort = sort_map[sort_by]\n",845 "\n",846 "print(f\"🌐 Target: VNDB Tag {tag_id} | Offset: {offset} | Fetching up to {num_vns} VNs (sorted by {sort_by})\")\n",847 "\n",848 "# Mount Google Drive\n",849 "print(\"🔄 Mounting Google Drive...\")\n",850 "drive.mount('/content/drive', force_remount=True)\n",851 "\n",852 "# Create working folder\n",853 "dataset_dir = \"/content/vndb_dataset\"\n",854 "os.makedirs(dataset_dir, exist_ok=True)\n",855 "\n",856 "# ====================== CALCULATE PAGINATION WITH OFFSET ======================\n",857 "start_page = (offset // MAX_RESULTS_PER_PAGE) + 1\n",858 "skip_in_first_page = offset % MAX_RESULTS_PER_PAGE\n",859 "\n",860 "print(f\"📌 Calculated start_page = {start_page}, skip first {skip_in_first_page} items on that page\")\n",861 "\n",862 "# ====================== FETCH VIA VNDB KANA API ======================\n",863 "pairs = []\n",864 "page = start_page\n",865 "items_collected = 0\n",866 "\n",867 "while len(pairs) < num_vns:\n",868 " payload = {\n",869 " \"filters\": [\"tag\", \"=\", tag_id],\n",870 " \"fields\": \"id, title, image.url\",\n",871 " \"sort\": selected_sort[\"sort\"],\n",872 " \"reverse\": selected_sort[\"reverse\"],\n",873 " \"results\": MAX_RESULTS_PER_PAGE,\n",874 " \"page\": page\n",875 " }\n",876 "\n",877 " # ==================== FULL DEBUG PRINT ====================\n",878 " print(f\"\\n📄 === API REQUEST PAGE {page} (offset={offset}) ===\")\n",879 " print(\"Payload sent:\")\n",880 " print(json.dumps(payload, indent=2))\n",881 " # ========================================================\n",882 "\n",883 " try:\n",884 " response = requests.post(API_URL, headers=HEADERS, json=payload, timeout=30)\n",885 "\n",886 " print(f\" 📡 Status code: {response.status_code}\")\n",887 "\n",888 " if response.status_code != 200:\n",889 " print(\" ❌ ERROR RESPONSE BODY:\")\n",890 " try:\n",891 " error_json = response.json()\n",892 " print(json.dumps(error_json, indent=2))\n",893 " except:\n",894 " print(response.text[:1000])\n",895 " break\n",896 "\n",897 " data = response.json()\n",898 " results = data.get(\"results\", [])\n",899 "\n",900 " if not results:\n",901 " print(\"✅ No more results.\")\n",902 " break\n",903 "\n",904 " # Handle offset skipping on the very first page we fetch\n",905 " if page == start_page and skip_in_first_page > 0:\n",906 " print(f\" ⏭️ Skipping first {skip_in_first_page} items due to offset\")\n",907 " results = results[skip_in_first_page:]\n",908 " skip_in_first_page = 0\n",909 "\n",910 " added = 0\n",911 " for vn in results:\n",912 " if len(pairs) >= num_vns:\n",913 " break\n",914 "\n",915 " title = vn.get(\"title\", \"\").strip()\n",916 " if not title:\n",917 " continue\n",918 "\n",919 " img_url = vn.get(\"image\", {}).get(\"url\") if isinstance(vn.get(\"image\"), dict) else None\n",920 "\n",921 " if img_url:\n",922 " pairs.append((title, img_url))\n",923 " added += 1\n",924 " items_collected += 1\n",925 "\n",926 " print(f\" → Added {added} VNs (total so far: {len(pairs)})\")\n",927 "\n",928 " if not data.get(\"more\", False):\n",929 " print(\"✅ Reached the end of results.\")\n",930 " break\n",931 "\n",932 " except Exception as e:\n",933 " print(f\"❌ Exception on API page {page}: {e}\")\n",934 " break\n",935 "\n",936 " page += 1\n",937 " time.sleep(DELAY_BETWEEN_PAGES)\n",938 "\n",939 "if not pairs:\n",940 " print(\"\\n❌ No visual novels found in this offset range. Check debug output above.\")\n",941 "else:\n",942 " print(f\"\\n✅ Collected {len(pairs)} visual novels (offset {offset}). Downloading images and creating TXT files...\")\n",943 "\n",944 " # ====================== DOWNLOAD ENUMERATED FILES ======================\n",945 " downloaded = 0\n",946 " for idx, (title, img_url) in enumerate(pairs, start=1):\n",947 " num_str = f\"{idx:04d}\"\n",948 " img_path = f\"{dataset_dir}/{num_str}.jpg\"\n",949 " txt_path = f\"{dataset_dir}/{num_str}.txt\"\n",950 "\n",951 " try:\n",952 " img_response = requests.get(img_url, headers=HEADERS, timeout=15)\n",953 " if img_response.status_code == 200:\n",954 " with open(img_path, \"wb\") as f:\n",955 " f.write(img_response.content)\n",956 " with open(txt_path, \"w\", encoding=\"utf-8\") as f:\n",957 " f.write(title)\n",958 " downloaded += 1\n",959 " if downloaded % 10 == 0 or downloaded == len(pairs):\n",960 " print(f\" ✅ Saved {num_str}.jpg + {num_str}.txt\")\n",961 " else:\n",962 " print(f\"⚠️ Failed to download image {num_str} (HTTP {img_response.status_code})\")\n",963 " except Exception as e:\n",964 " print(f\"❌ Error downloading {num_str}: {e}\")\n",965 "\n",966 " # ====================== WRITE METADATA (index/offset/tag) ======================\n",967 " timestamp = datetime.datetime.now().strftime(\"%Y%m%d_%H%M%S\")\n",968 " with open(f\"{dataset_dir}/INFO.txt\", \"w\", encoding=\"utf-8\") as f:\n",969 " f.write(f\"VNDB Tag ID : {tag_id}\\n\")\n",970 " f.write(f\"Offset : {offset}\\n\")\n",971 " f.write(f\"Batch Size : {num_vns}\\n\")\n",972 " f.write(f\"Actual Downloaded: {downloaded}\\n\")\n",973 " f.write(f\"Sort Order : {sort_by}\\n\")\n",974 " f.write(f\"Start Page : {start_page}\\n\")\n",975 " f.write(f\"Collected on : {timestamp}\\n\")\n",976 " f.write(f\"File index 0001 = VN #{offset + 1} in the full tag list\\n\")\n",977 "\n",978 " print(\"📝 Metadata INFO.txt written (contains index/offset/tag info)\")\n",979 "\n",980 " # ====================== CREATE ZIP & SAVE TO DRIVE ======================\n",981 " zip_name = f\"vndb_{tag_id}_offset{offset:04d}_{num_vns}vns_{timestamp}\"\n",982 " zip_path_local = f\"/content/{zip_name}.zip\"\n",983 "\n",984 " print(f\"\\n🗜️ Creating ZIP file with {downloaded} image+txt pairs + INFO.txt...\")\n",985 " shutil.make_archive(f\"/content/{zip_name}\", 'zip', dataset_dir)\n",986 "\n",987 " drive_folder = \"/content/drive/MyDrive/vndb_datasets\"\n",988 " os.makedirs(drive_folder, exist_ok=True)\n",989 " drive_zip_path = f\"{drive_folder}/{zip_name}.zip\"\n",990 " shutil.copy(zip_path_local, drive_zip_path)\n",991 "\n",992 " print(\"\\n\" + \"=\"*80)\n",993 " print(\"🎉 SUCCESS! Your dataset is ready\")\n",994 " print(f\"📦 ZIP file: {zip_name}.zip\")\n",995 " print(f\"📤 Saved to Google Drive → {drive_zip_path}\")\n",996 " print(f\" • Contains: 0001.jpg + 0001.txt ... + INFO.txt (with offset/tag/index)\")\n",997 " print(f\" • Total pairs: {downloaded}\")\n",998 " print(\"=\"*80)\n",999 "\n",1000 " # ====================== PREVIEW FIRST 3 PAIRS ======================\n",1001 " print(\"\\n📸 Preview of first 3 pairs (click images to enlarge):\")\n",1002 " for i in range(min(3, len(pairs))):\n",1003 " num_str = f\"{i+1:04d}\"\n",1004 " img_file = f\"{dataset_dir}/{num_str}.jpg\"\n",1005 " display(HTML(f\"<h4>{num_str}. {pairs[i][0]}</h4>\"))\n",1006 " display(Image(filename=img_file, width=400))\n",1007 " print(\"─\" * 70)\n",1008 "\n",1009 " print(f\"\\n✅ All files are also in: {dataset_dir} (you can download the folder manually if needed)\")"1010 ]1011 },1012 {1013 "cell_type": "markdown",1014 "source": [1015 "# 📌 Pinterest fetch"1016 ],1017 "metadata": {1018 "id": "JrtsI98cmAxB"1019 }1020 },1021 {1022 "cell_type": "markdown",1023 "metadata": {1024 "id": "HO3NmF03QDpt"1025 },1026 "source": [1027 "Pinterest board downloader\n",1028 "\n",1029 "⚠️️ Use a throwaway account! Pinterest is a stupid website run by AI bots.\n",1030 "\n",1031 "1. Install the EditThisCookie (or \"Get cookies.txt LOCALLY\") Chrome extension.\n",1032 "2. Log into Pinterest in Chrome → open your board.\n",1033 "3. Click the extension icon → Export cookies for pinterest.com → save as cookies.txt (plain text / Netscape format).\n",1034 "4. In Colab, click the folder icon (left sidebar) → upload cookies.txt to your google drive."1035 ]1036 },1037 {1038 "cell_type": "code",1039 "source": [1040 "# ==================== SINGLE CELL - FULL PINTEREST BOARD DOWNLOADER (Cookies from Google Drive) ====================\n",1041 "\n",1042 "# Install gallery-dl\n",1043 "!pip install -q gallery-dl\n",1044 "\n",1045 "import os\n",1046 "from google.colab import files\n",1047 "from google.colab import drive\n",1048 "import shutil # Import shutil for copying files\n",1049 "\n",1050 "# ====================== CONFIGURATION ======================\n",1051 "# 1. Paste your board URL here\n",1052 "board_url = \"\" #@param {type:\"string\"}\n",1053 "\n",1054 "# 2. Path to your cookies file on Google Drive (change only if it's in a subfolder)\n",1055 "cookies_file = \"/content/drive/MyDrive/pinterest_cookies.txt\" #@param {type:\"string\"}\n",1056 "\n",1057 "# Optional: custom board name (auto-detected from URL by default)\n",1058 "board_name = board_url.rstrip(\"/\").split(\"/\")[-1]\n",1059 "#or \"pinterest_board\" #@param {type:\"string\"}\n",1060 "\n",1061 "print(\"✅ Board URL:\", board_url)\n",1062 "print(\"📁 Board name:\", board_name)\n",1063 "print(\"🔑 Cookies path:\", cookies_file)\n",1064 "\n",1065 "# ====================== MOUNT GOOGLE DRIVE ======================\n",1066 "print(\"🚀 Mounting Google Drive...\")\n",1067 "drive.mount('/content/drive', force_remount=False)\n",1068 "\n",1069 "# Check if cookies file exists\n",1070 "if os.path.exists(cookies_file):\n",1071 " print(\"✅ Cookies file found! Full board (700+ images) will be downloaded.\")\n",1072 "else:\n",1073 " print(\"❌ Cookies file NOT found at the path above. Only ~200 images will download.\")\n",1074 "\n",1075 "# ====================== CREATE OUTPUT FOLDER ======================\n",1076 "output_dir = f\"/content/{board_name}\"\n",1077 "os.makedirs(output_dir, exist_ok=True)\n",1078 "\n",1079 "print(\"🚀 Starting download... (this can take a while for large boards)\")\n",1080 "\n",1081 "# ====================== BUILD & RUN GALLERY-DL COMMAND ======================\n",1082 "cmd = f'gallery-dl --dest \"{output_dir}\"'\n",1083 "if os.path.exists(cookies_file):\n",1084 " cmd += f' --cookies \"{cookies_file}\"'\n",1085 "cmd += f' \"{board_url}\"'\n",1086 "\n",1087 "# Execute the download\n",1088 "!{cmd}\n",1089 "\n",1090 "# Count downloaded files\n",1091 "total_files = sum([len(files) for r, d, files in os.walk(output_dir)])\n",1092 "print(f\"✅ Download finished! {total_files} files saved in {output_dir}\")\n",1093 "\n",1094 "# ====================== ZIP & AUTO-DOWNLOAD ======================\n",1095 "zip_path = f\"/content/{board_name}.zip\"\n",1096 "print(\"📦 Zipping all images...\")\n",1097 "!zip -r -q \"{zip_path}\" \"{output_dir}\"\n",1098 "\n",1099 "print(f\"✅ Zipped everything → {zip_path}\")\n",1100 "\n",1101 "# Save to Google Drive\n",1102 "drive_zip_destination = f\"/content/drive/MyDrive/{board_name}.zip\"\n",1103 "shutil.copy2(zip_path, drive_zip_destination)\n",1104 "print(f\"✅ Copied '{zip_path}' to Google Drive at '{drive_zip_destination}'\")\n",1105 "\n",1106 "# Auto-download the zip to your computer\n",1107 "#files.download(zip_path)\n",1108 "\n",1109 "#print(\"🎉 All done! Your full Pinterest board is now downloaded, zipped, and available in your Google Drive and local downloads.\")"1110 ],1111 "metadata": {1112 "id": "E7vSr64JkWWa",1113 "cellView": "form"1114 },1115 "execution_count": null,1116 "outputs": []1117 },1118 {1119 "cell_type": "code",1120 "execution_count": null,1121 "metadata": {1122 "id": "g0524iIvU1I_"1123 },1124 "outputs": [],1125 "source": [1126 "# Auto-download the zip file\n",1127 "files.download(zip_path)"1128 ]1129 },1130 {1131 "cell_type": "markdown",1132 "source": [1133 "# 🖼️ Fetch from Sankaku Complex"1134 ],1135 "metadata": {1136 "id": "s40baE0G_wn7"1137 }1138 },1139 {1140 "cell_type": "code",1141 "source": [1142 "from google.colab import drive, userdata\n",1143 "\n",1144 "# Mount your Google Drive\n",1145 "drive.mount('/content/drive')\n",1146 "\n",1147 "#@markdown Load Sankaku Complex secrets (set these in Colab's left sidebar → Secrets panel first!)\n",1148 "SANKAKU_USERNAME = userdata.get('SANKAKU_USERNAME')\n",1149 "SANKAKU_PASSWORD = userdata.get('SANKAKU_PASSWORD')\n",1150 "\n",1151 "print(\"✅ Google Drive mounted and Sankaku credentials loaded from secrets!\")"1152 ],1153 "metadata": {1154 "cellView": "form",1155 "id": "XmDavSOowJSJ"1156 },1157 "execution_count": null,1158 "outputs": []1159 },1160 {1161 "cell_type": "code",1162 "source": [1163 "# @title Sankaku Advanced Downloader + Pagination + Fullsize Processing\n",1164 "\n",1165 "search_phrase = \"\" # @param {type:\"string\"}\n",1166 "N = 291 # @param {type:\"slider\", min:0, max:2000, step:100}\n",1167 "offset = 0 # @param {type:\"slider\", min:0, max:20000, step:10}\n",1168 "\n",1169 "download_thumbnails = True # @param {type:\"boolean\"}\n",1170 "create_image_text_pairs = True # @param {type:\"boolean\"}\n",1171 "download_fullsize = True # @param {type:\"boolean\"}\n",1172 "list_animated_links = True # @param {type:\"boolean\"}\n",1173 "download_animated_and_extract_frames = False # @param {type:\"boolean\"}\n",1174 "\n",1175 "# ── NEW: Fullsize post-processing for training datasets ──\n",1176 "process_fullsize_for_training = False # @param {type:\"boolean\"}\n",1177 "target_shortest_side = 1024 # @param {type:\"slider\", min:512, max:2048, step:64}\n",1178 "output_drive_folder_name = \"sankaku_processed\" # @param {type:\"string\"}\n",1179 "output_zip_name = \"resized_image_text_pairs.zip\" # @param {type:\"string\"}\n",1180 "\n",1181 "import requests\n",1182 "import os\n",1183 "import shutil\n",1184 "from zipfile import ZipFile\n",1185 "import urllib.parse\n",1186 "import subprocess\n",1187 "import time\n",1188 "from http.cookiejar import MozillaCookieJar\n",1189 "from PIL import Image\n",1190 "from tqdm.notebook import tqdm\n",1191 "from pathlib import Path\n",1192 "import gc\n",1193 "\n",1194 "# ====================== LOGIN ======================\n",1195 "login_url = \"https://sankakuapi.com/auth/token\"\n",1196 "payload = {\"login\": SANKAKU_USERNAME, \"password\": SANKAKU_PASSWORD}\n",1197 "headers = {\n",1198 " \"Accept\": \"application/vnd.sankaku.api+json;v=2\",\n",1199 " \"Origin\": \"https://sankaku.app\",\n",1200 " \"User-Agent\": \"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36\",\n",