Team Ai
Datasetpublic

Puzer/github-repo-embeddings

GitHub Repo Embeddings (Dataset) This dataset contains: GitHub repository embeddings learned from star co-occurrence. Raw data for training such embeddings (2016 - 2025 years) It is generated by the same pipeline as this repo and is intended for offline analysis, research, and downstream search/indexing. See Demo which uses trained embeddings Summary Source: GitHub Archive (BigQuery) WatchEvent + repo metadata. Signal: repositories starred together by the same… See the full description on the dataset page: https://huggingface.co/datasets/Puzer/github-repo-embeddings.

sourceHugging Facecc-by-4.0updated 9mo agoView on Hugging Face
6likes51downloads
Dataset Card

GitHub Repo Embeddings (Dataset)

This dataset contains:

  • —GitHub repository embeddings learned from star co-occurrence.
  • —Raw data for training such embeddings (2016 - 2025 years)

It is generated by the same pipeline as this repo and is intended for offline analysis, research, and downstream search/indexing.

See Demo which uses trained embeddings

Summary

  • —Source: GitHub Archive (BigQuery) WatchEvent + repo metadata.
  • —Signal: repositories starred together by the same user.
  • —Model: torch.nn.EmbeddingBag trained with MultiSimilarityLoss.
  • —Embedding size: 128 dims.

Files

starred_repos.parquet

User-level training data.

  • —repo_ids: list[int], repo ids starred by a user (order preserved from events).

repos_meta.parquet

Repository metadata aligned with the training data.

  • —repo_id: int
  • —repo_name: str (owner/name)
  • —stars: int, frequency of stars in this dataset
  • —created_at: datetime, repo creation date (first push event)
  • —last_updated: datetime, last push event

repo_embeddings_with_meta.parquet

Repository metadata + learned embeddings aligned by repo_id.

  • —Includes columns from repos_meta.parquet
  • —embedding: list[float], 128-dim vector

Notes

  • —The dataset is derived from public GitHub Archive data and is intended for research and demo purposes.