Ethical approval and patient sample collection
All ethical approvals required for patient sample collection, organoid derivation, and the academic and commercial use of derived organoid models were obtained from participating clinical sites and the Wellcome Sanger Institute. The study was approved under IRAS ID 203519 and REC reference 16/LO/1110. Additional information is available in the Supplementary Information.
Organoid biobank availability
The organoid models received ethical approval for academic and commercial applications. Models are distributed by the American Type Culture Collection (ATCC) through the HCMI collection (https://www.atcc.org/hcmi) and by EMD Millipore. Repository information and model availability are listed in Supplementary Table 2.
Organoid derivation from primary tumour tissue
Comprehensive protocols for organoid derivation, cryopreservation and routine culture are available on protocols.io. These protocols include workflow diagrams and instructional videos62,63,64.
Tumour tissue was washed three times with PBS, minced, and either cryopreserved after centrifugation at 800g for 2 min63 or enzymatically digested for 1–2 h at 37 °C. The resulting suspension was passed through a 100 μm filter, centrifuged and washed to remove cellular debris and digestion buffer62. Isolated cells were embedded in approximately 15 μl droplets of extracellular matrix comprising 80:20 basement membrane extract (BME):medium, with a final protein concentration of 6.4–9.6 mg ml−1 (Cultrex BME Type 2 Select 3532-001-02). Droplets were plated in pre-warmed 6-well plates according to established methods64.
Following matrix polymerization for 15–20 min at 37 °C, 2 ml of organoid culture medium prepared using published formulations39,65,66,67 was added. The medium was supplemented with antibiotics and 1 μl ml−1 ROCK inhibitor Y-27632 (Staratech Scientific, S1049-SEL-5mg).
Once cultures expanded to at least 25 million cells, 25 cryovials were prepared for biobanking and cell pellets were collected for sequencing. Post-thaw viability was assessed by re-culturing organoids for four passages and performing freeze–thaw quality-control testing.
Organoid culture and passaging
Organoids were maintained either in 80% BME-2 droplets or in 5% BME-2 suspension culture38. The 80% BME-2 cultures were maintained as described above. For suspension culture, cancer organoids were grown in a diluted medium–extracellular matrix (ECM) mixture. For example, 10 ml of organoid medium was combined with 500 μl of BME-2 to produce the required concentration. Organoids were then cultured in ultra-low-attachment flasks or plates.
For passaging, the culture medium, organoids and ECM were collected and centrifuged at 800g for 2 min. After removing the supernatant, organoids were dissociated with TrypLE (Gibco, 12604013). Samples were incubated for 10–60 min until the organoids formed small cell clumps. Cells were pelleted and replated using the established protocol64.
Organoid model identifiers
Each model described in this study has a Sanger ID beginning with “WTSI”, referring to the Wellcome Trust Sanger Institute. Models shared with the HCMI collection also have an HCMI identifier beginning with “HCM”. Where both identifiers are available, the HCMI nomenclature is used as the sample ID. Identifier equivalences are provided in Supplementary Table 2.
Whole-genome sequencing and genomic analysis
Whole-genome sequencing and read alignment
DNA was extracted from snap-frozen tumour tissue, snap-frozen organoid cell pellets, blood or formalin-fixed paraffin-embedded normal tissue for whole-genome sequencing (WGS). Paired-end sequencing reads of 150 bp were generated on the Illumina HiSeq X Ten platform at an average coverage of 38×, comparable to the coverage used by PCAWG. Reads were aligned with BWA-MEM v0.7.1768. PCR duplicates, unmapped reads and non-uniquely mapped reads were removed before downstream analysis.
Somatic SNV and indel detection
Somatic single-nucleotide variants (SNVs) and short insertions and deletions (indels) were identified using cgpCaVEMan69 and cgpPindel70, respectively. Germline variants and technical artefacts were filtered using matched normal samples and a panel of normal samples. Additional processing was performed with cgpCaVEManPostProcessing (https://github.com/cancerit). Variant allele frequencies (VAFs) were estimated with vafCorrect v2.4.071, and variants with VAF > 0.05 were retained. Further methodological details are provided in the Supplementary Information.
Structural variant and copy-number analysis
Copy-number and allele-specific profiles were generated using AMBER v3.5 and COBALT v1.1172. Somatic and germline SNVs and indels were identified with SAGE v2.8 and annotated using SnpEff v5.073. Somatic structural variants (SVs) were detected with GRIDSS2 v2.12.074, annotated using RepeatMasker v4.1.275 and kraken2 v2.1.276, and filtered with GRIPSS v1.977.
These data were integrated with PURPLE v2.5472 to estimate microsatellite status, tumour purity, ploidy, whole-genome doubling (WGD) and somatic copy-number alterations (SCNAs). SCNAs were subsequently converted to discrete copy-number states. SAGE, GRIPSS, AMBER, COBALT and PURPLE were developed by the Hartwig Medical Foundation (HMF) (https://github.com/hartwigmedical/hmftools). Additional details are available in the Supplementary Information.
Cancer driver gene and mutation annotation
We generated a panel of 783 cancer driver genes by combining two complementary gene sets from the IntOGen78 and COSMIC79 databases, as previously described9. Each gene was classified according to its mechanism of action as activating (Act; oncogenes), loss-of-function (LoF; tumour suppressor genes), ambiguous, when evidence supported both mechanisms, or fusion. The full cancer driver gene list is available at https://cellmodelpassports.sanger.ac.uk/downloads.
Putative cancer driver mutations, including frameshift, nonsense, stop-lost, exonic splicing silencer, missense and in-frame variants, were collected from SNVs and indels affecting these genes. Annotations were obtained from IntOGen, including Cancer Genome Interpreter and BoostDM78, MSKCC80, and cancer-predisposition variant datasets. Predisposition variants were identified by overlap with a reference set of pathogenic germline variants with matching effects81.
Loss of heterozygosity
A tumour suppressor or ambiguous gene was classified as having loss of heterozygosity when the minor-allele DNA copy number was <0.5 and the difference between rounded ploidy and rounded total copy number was >0, indicating no amplification of the non-mutated allele. Loss of heterozygosity was also assigned when the VAF of a loss-of-function mutation was >0.85.
Biallelic genomic alterations
Biallelic alterations included all loss-of-heterozygosity events, homozygous deletions, structural-variant disruptions and cases in which each allele was affected by a separate loss-of-function mutation.
Mutation multiplicity, cancer cell fraction and clonal variants
SNVs were intersected with SCNA segments using the GRanges and findOverlaps functions from the GenomicRanges R package v1.56.1. This provided the major-allele, minor-allele and total tumour copy numbers for each SNV, together with VAF and tumour-purity estimates. Mutation multiplicity, defined as the number of chromosomal copies carrying a mutation, and cancer cell fraction (CCF), defined as the proportion of cancer cells carrying the alteration, were calculated using the formulas reported by Steele et al.82 and Dentro et al.83. Mutations with CCF > 0.75 were classified as clonal.
Copy-number correlation between paired organoids and tumours
SCNA segment data were divided into 100 kb genomic bins. The copy number for each bin was calculated as the mean copy number across all positions within the 100,000 bp window. Regions included in the ENCODE blacklist84 (https://github.com/Boyle-Lab/Blacklist) were excluded. Pearson correlation coefficients were calculated by comparing mean copy numbers between paired tumour and organoid samples.
Focal amplification criteria
Focal amplification was defined as at least two genomic segments, each larger than 100 kb, with a log2(CN/ploidy) value > 6.
Mutational signature analysis
Mutational signatures were extracted separately for tumour and organoid samples within each tumour type using SigProfilerExtractor v1.2.285. COAD/READ microsatellite-instability-high (MSI) samples were analysed separately from COAD/READ microsatellite-stable (MSS) samples. STAD samples were excluded because of the limited sample size.
The resulting mutation matrix was analysed with SigProfilerAssignment v0.2.530 to assign known COSMIC v3.4 single-base substitution (SBS) signatures. De novo signatures were refitted using COSMIC v3.4 and an additional set of novel colorectal cancer and MSI-specific signatures identified and validated by the Mutographs project86.
For visualization, signatures accounting for less than 10% of mutations in an individual sample were grouped as “Others”. These signatures were excluded when calculating the proportion of models containing each signature and the proportion of mutations attributed to each signature by cancer type. SBS5 and SBS40a were combined into one category based on the hypothesis that they reflect correlated mutational processes87,88.
Homologous recombination deficiency
Homologous recombination deficiency was evaluated using CHORD v2.0.334.
Complex genomic rearrangements
Chromothripsis and other complex genomic rearrangements were identified with ShatterSeek v1.126, following previously published methods89.
RNA sequencing and transcriptomic analysis
Paired-end transcriptome reads of 75 bp were quality-filtered and aligned to GRCh38, Ensembl build 98, using STAR v2.5.0c90 with standard parameters (https://github.com/cancerit/cgpRna). BAM files were processed with RSEM v1.3.391 to obtain per-gene read counts and transcripts-per-million (TPM) values. TPM data were used for downstream analyses.
CMS and CRIS molecular subtypes
Consensus molecular subtypes (CMS) and colorectal cancer intrinsic subtypes (CRIS) were inferred for COAD/READ organoids with CMSCaller36,92 using expected gene-count data generated by RSEM.
CRISPR–Cas9 genetic screening
Detailed protocols for generating Cas9-expressing organoids and transducing organoid cultures with CRISPR–Cas9 libraries are available on protocols.io93,94. Stable Cas9-expressing organoids were generated using lentiCas9-Blast (Addgene 52962) and polybrene at 8 μg ml−1. After overnight incubation, the medium was replaced with complete medium containing 2.5 μM Y-27632. Blasticidin selection (Invivogen, ant-bl-1, 10 mg ml−1) began 120 h after transduction.
Cas9 activity was measured using a fluorescent reporter assay (Addgene 67982 and 67981)95. Only organoid lines with at least 75% Cas9 activity were advanced to single-guide RNA (sgRNA) library transduction.
Two genome-wide CRISPR–Cas9 sgRNA libraries were used: the Human CRISPR Library Yusa v1.1 (ref. 3), containing 100,086 sgRNAs targeting 18,009 genes, with 5–10 sgRNAs per gene and 1,004 non-targeting controls; and the Minimal Genome-Wide Human CRISPR–Cas9 Library (MinLibCas9, Addgene 164896)41. MinLibCas9 contains 37,522 sgRNAs targeting 18,761 genes, with two optimized sgRNAs per gene and 200 non-targeting controls (Extended Data Fig. 8a).
The lentiviral volume required to achieve a multiplicity of infection of 0.3 was determined by titration. Transduction efficiency was measured using BFP flow cytometry. A total of 12.5 × 108 cells for Yusa v1.1 or 3.3 × 107 cells for MinLibCas9 were transduced in triplicate using the same batch of Cas9-expressing organoids. No significant batch effects were observed between independent Cas9-organoid batches (Extended Data Fig. 8b).
Cells were infected with lentiviral whole-genome sgRNA libraries at 100× coverage in medium containing 8 μg ml−1 polybrene and 2.5 μM Y-27632. After overnight incubation, cells were plated in fresh medium as a 5% suspension culture. A target transduction efficiency of 30% was confirmed on day 6. Screens were maintained under puromycin selection (Invivogen, ant-pr-1, 10 mg ml−1) for an additional 16 days, giving a total screening period of 21 days. Screens required a final selection efficiency of at least 60% to pass quality control. Approximately 2.5 × 107 cells were collected at the end of each screen, pelleted and stored at −80 °C for downstream processing.
CRISPR screen data processing
CRISPR screens generated with the Yusa v1.1 (ref. 3) and MinLibCas941 libraries were harmonized by restricting analyses to sgRNAs shared between the two libraries. Quality-control procedures adapted from established cell-line screening pipelines3 assessed replicate concordance, sgRNA representation and the ability to distinguish essential from non-essential genes.
Low-quality replicates and organoid models that failed predefined quality criteria were excluded. Read counts were normalized and adjusted for copy-number effects with CRISPRcleanR v3.0.196, followed by batch correction across libraries97. Gene-level fitness effects were estimated using BAGEL v245 and curated essential-gene reference sets. Log-fold-change (LFC) values were scaled against essential and non-essential controls to support comparisons across organoid models. Complete information on sgRNA selection, quality thresholds, normalization, batch correction, statistical analysis and final model selection is provided in the Supplementary Information.
Processed CRISPR screen analysis
Core fitness genes in organoids
ADaM, implemented in the CoRe R package v1.0.0 (https://github.com/DepMap-Analytics/CoRe)98, was applied with the previously defined curated essential-gene reference set3 to calculate false-positive rates. This analysis identified 751 core essential genes (Supplementary Table 4). Although all organoids were analysed together, only three cancer types were included; PAAD and STAD organoids were excluded. Therefore, these core fitness genes should not be interpreted as a pan-cancer set.
Comparison with core fitness genes identified in cancer cell lines9 showed an overlap of 654 genes and identified 97 organoid-specific genes (Extended Data Fig. 9). Organoid-specific core fitness genes were analysed using Fisher’s exact test for pathway enrichment across Gene Ontology Biological Processes, KEGG pathways99 and Hallmark gene sets100 obtained from the msigdbr R package v7.5.1101. PanCancer common essential genes identified using the AUC method were excluded98. Where several pathways contained the same organoid-specific core fitness genes, only the pathway with the most significant P value was shown.
Differential dependency analysis in gastrointestinal, COAD/READ and ESCA organoids
We identified genes with variable dependency across organoid models within each cancer type to characterize context-specific genetic dependencies. The analysis followed three filtering steps:
-
1.
Gene filtering based on consistent depletion or non-depletion. Binary dependency matrices generated with BAGEL2 for each cancer type were used to exclude genes classified as depleted in only one organoid or non-depleted in only one organoid.
-
2.
Removal of core fitness and control genes. Genes identified as pan-cancer core fitness genes in cell-line datasets, the 751 organoid-specific core fitness genes identified here, and essential and non-essential control genes were excluded to focus on differential dependency patterns.
-
3.
Expression filtering. Genes with a mean expression level of log2(TPM + 1) < 0.1 within a cancer type were considered not expressed and excluded.
For each gene that passed these filters, we calculated the difference in LFC between depleted and non-depleted organoids and assessed statistical significance using Fisher’s exact test. Genes with an FDR-adjusted P value < 0.05 were classified as differentially dependent. This identified 7,086 genes in gastrointestinal organoids, 5,103 genes in COAD/READ organoids and 4,082 genes in ESCA organoids (Supplementary Table 7).
Biomarker analysis of gene dependencies
We performed a systematic biomarker analysis across gastrointestinal, COAD/READ and ESCA organoids to identify molecular and clinical features associated with context-specific gene dependencies. Candidate dependencies were selected from the differential dependency analysis according to recurrence criteria and tested against curated genomic, transcriptomic and clinical features.
Analysed features included driver mutations and specific variants, copy-number alterations, structural variants, mutational signatures, gene expression, pathway activity scores, and composite loss-of-function and gain-of-function events. Associations between gene fitness effects and candidate biomarkers were evaluated using linear regression models incorporating relevant technical and biological covariates. Statistical significance was assessed using likelihood-ratio tests and multiple-testing correction.
Significant associations were prioritized according to adjusted P value and effect size and assigned to tiers reflecting the strength of evidence. Full details of feature selection, model design, statistical thresholds and classification criteria are provided in the Supplementary Information.
High-throughput organoid drug screening
For high-throughput drug screening, organoids were dissociated into single cells, counted and seeded at model-specific optimized densities in 5% BME-2 suspension cultures. After 96 h, allowing organoids to reform, assay plates were prepared with a 50% BME-2:organoid medium base layer. Organoids were transferred to 384-well plates using Multidrop Combi dispensers (Thermo Scientific).
Compounds were added 24 h later using an Echo555 acoustic dispenser (Labcyte), and cells were treated for 72 h. Viability was measured with CellTiter-Glo 2.0 (Promega).
Two independent screening projects were performed. The first used a 7 × 7 concentration matrix, providing 49 measurements for each drug combination. The second used a reduced 25-point matrix covering the same concentration range. Compounds were tested in biological duplicate in the first project and as single replicates in the second. Agents included in multiple combinations, such as afatinib, received additional technical replication and were also tested as monotherapies.
Raw viability measurements were analysed separately for each project. Data were normalized by plate using negative controls, including untreated, DMSO and blank wells, and positive controls, including MG-132 and staurosporine. Dose-response curves were fitted with a non-linear mixed-effects model to estimate IC50 and area-under-the-curve (AUC) values using the gdscIC50 R package v1.7.3102.
For drug combinations, maximum combination effect (combo_MaxE) was defined as the second-highest measured inhibition. Bliss excess was calculated as the observed combination inhibition minus the inhibition predicted by Bliss additivity from the corresponding monotherapies. Reported values represent the mean across both screening projects.
Validation of drug sensitivity
For validation experiments, organoids were dissociated into single cells, resuspended in organoid medium and seeded into 96-well plates over a 50% BME-2:organoid medium base layer. Depending on the model, 2,000–5,000 cells were plated per well.
For monotherapy experiments, nine drug concentrations spanning a 256-fold range were added four days after plating, with technical triplicates for each condition. Combination experiments tested five concentrations of the KRAS inhibitors sotorasib or MRTX1133 across a 256-fold range together with two fixed concentrations of the EGFR and ERBB2 inhibitors afatinib and gefitinib. All combination conditions were tested in triplicate.
Cell viability was measured with CellTiter-Glo 2.0 at 72 h after monotherapy treatment and at 0, 3, 6 and 9 days after combination treatment. Dose-response curves were fitted using GraphPad Prism. Each condition included three technical replicates and two biological replicates.
EGF depletion experiments
Organoids were dissociated into single cells and resuspended in organoid culture medium that was either supplemented with EGF or deprived of EGF. Cells were seeded into 96-well plates containing a 50:50 medium-to-ECM base layer. The bottom ECM-containing layer was prepared using the same EGF condition as the overlaying medium. Cell viability was measured with CellTiter-Glo 2.0 seven days after seeding.
Reporting summary
Additional information about the study design and reporting is available in the Nature Portfolio Reporting Summary linked to this article.
Source: www.nature.com


