Team Ai
Apppublic

JacksonFW/transcriptomics-explorer

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes
README.md164 linesDownload Raw Back to analysis
1# Analysis Pipeline2 3R scripts to prepare RNA-seq count data and run differential expression analysis. The output CSV can be uploaded directly to the Transcriptomics Explorer dashboard.4 5---6 7## Workflow8 9```10fetch_geo.R  (Option A: airway built-in dataset, OR Option B: any GEO accession)11    ↓12counts.csv + metadata.csv13    ↓14deseq2_analysis.R  (or edger_analysis.R)15    ↓16de_results.csv  (with gene symbols)17    ↓18Upload to dashboard → all 6 tabs update instantly19```20 21---22 23## Files in this folder24 25| File | Purpose |26|------|---------|27| `fetch_geo.R` | Get a count matrix — either from the built-in airway dataset or from GEO |28| `deseq2_analysis.R` | Run DESeq2 → outputs `de_results.csv` with gene symbols |29| `edger_analysis.R` | Alternative to DESeq2 (better for < 3 replicates per group) |30| `counts.csv` | Count matrix generated by fetch_geo.R (airway dataset) |31| `metadata.csv` | Sample info generated by fetch_geo.R |32| `de_results.csv` | Final output — upload this to the dashboard |33 34---35 36## Requirements37 38**R ≥ 4.1** is required. All packages install automatically on first run.39 40Packages installed automatically:41- `airway` — built-in Bioconductor RNA-seq teaching dataset (Option A)42- `GEOquery` — download datasets from GEO (Option B)43- `DESeq2` — differential expression (recommended)44- `edgeR` + `limma` — alternative DE method45- `org.Hs.eg.db` + `AnnotationDbi` — gene symbol annotation (human)46- `ashr` — fold change shrinkage47- `ggplot2`, `dplyr` — plotting and data manipulation48 49For **mouse** data, also install: `BiocManager::install("org.Mm.eg.db")`50 51---52 53## Step 1 — Get count matrix54 55```r56# In RStudio: open fetch_geo.R and click Run57# In Terminal:58Rscript fetch_geo.R59```60 61`fetch_geo.R` has two options at the top — set `USE_AIRWAY`:62 63```r64USE_AIRWAY <- TRUE    # TRUE = airway built-in (no download, recommended for testing)65USE_AIRWAY <- FALSE   # FALSE = download from GEO using GEO_ID below66GEO_ID     <- "GSE96870"67```68 69**Option A — airway (default, recommended for testing):**70- No internet download required71- 8 samples, ~16,000 genes after filtering72- Experiment: dexamethasone drug treatment vs untreated, human airway cells73- Produces the `de_results.csv` included in this folder74 75**Option B — GEO download:**76- Downloads any GSE accession that provides a supplementary count matrix file77- NOTE: after downloading, open `metadata.csv` to find the correct column names for Step 278 79---80 81## Step 2 — Run differential expression82 83### Option A: DESeq2 (recommended, ≥ 3 replicates per group)84 85```r86Rscript deseq2_analysis.R87```88 89### Option B: edgeR (works with fewer replicates)90 91```r92Rscript edger_analysis.R93```94 95**Edit the USER SETTINGS at the top of the script before running:**96 97For the **airway dataset** (default):98```r99SAMPLE_ID_COL <- "sample_id"100CONDITION_COL <- "dex"101GROUP_A       <- "untrt"    # reference / control102GROUP_B       <- "trt"      # treatment / case103GENE_ID_TYPE  <- "ENSEMBL"  # airway uses Ensembl IDs104ANNOTATION_DB <- org.Hs.eg.db  # human annotation105```106 107For a **GEO dataset**, first inspect `metadata.csv`:108```r109meta <- read.csv("metadata.csv")110colnames(meta)                       # find the condition column name111table(meta$characteristics_ch1)      # see exact group label text112```113Then set `SAMPLE_ID_COL`, `CONDITION_COL`, `GROUP_A`, `GROUP_B` accordingly.114 115**Gene annotation:**116The scripts automatically map Ensembl IDs (`ENSG00000...`) to gene symbols (`TP53`, `BRCA1`) using `org.Hs.eg.db`. Genes that cannot be mapped keep their Ensembl ID. For mouse data, change `ANNOTATION_DB <- org.Mm.eg.db`.117 118**Output files:**119| File | Description |120|------|-------------|121| `de_results.csv` | DESeq2/edgeR results with gene symbols → upload to dashboard |122| `ma_plot.pdf` | MA plot QC check |123| `volcano_plot.pdf` | Volcano plot preview |124 125---126 127## Step 3 — Upload to dashboard128 1291. Start the dashboard from the parent folder:130   ```bash131   python app.py132   ```1332. Open **http://localhost:7860** in your browser1343. Click **Upload** and select `de_results.csv`1354. All 6 tabs update automatically136 137---138 139## Column format of de_results.csv140 141| Column | Description |142|--------|-------------|143| `gene` | Gene symbol (e.g. TP53) — or Ensembl ID if symbol not found |144| `ensembl_id` | Original Ensembl ID (e.g. ENSG00000141510) |145| `baseMean` | Average expression across all samples |146| `log2FoldChange` | Log2 ratio of GROUP_B / GROUP_A |147| `pvalue` | Raw p-value |148| `padj` | Adjusted p-value (Benjamini-Hochberg FDR) |149| `regulation` | "Up", "Down", or "NS" based on thresholds |150 151---152 153## Choosing between DESeq2 and edgeR154 155| | DESeq2 | edgeR |156|-|--------|-------|157| Best for | ≥ 3 replicates per group | 2–3 replicates, smaller datasets |158| Normalisation | Median of ratios | TMM |159| Fold change | Shrinkage via ashr (more conservative) | Raw logFC |160| Gene annotation | Yes — via org.Hs.eg.db | Yes — via org.Hs.eg.db |161| Speed | Slightly slower | Slightly faster |162 163Both produce equivalent results on well-powered experiments. DESeq2 is the current community default.164