Genetic Analysis of the Triglyceride-to-HDL Cholesterol Ratio: Study Population and Methods
SEO keywords: TG:HDL ratio, triglyceride-to-HDL cholesterol ratio, genetic association study, cardiometabolic disease, liver disease, metabolic health, exome sequencing
This comprehensive study investigated the genetic, clinical, imaging and molecular factors associated with the triglyceride-to-high-density lipoprotein cholesterol ratio (TG:HDL ratio). Researchers combined data from more than one million participants across population-based and health-system cohorts, integrating epidemiological analyses, genome-wide association studies, exome sequencing, liver imaging, transcriptomics, proteomics and experimental mouse models.
Study Population
The genetic discovery analysis of the TG:HDL ratio included 1,032,116 participants recruited from population-based and hospital-based health systems across three continents. The participating cohorts were:
- UK Biobank (UKB): 409,602 participants
- Geisinger Health System MyCode cohort (GHS): 143,129 participants
- Mexico City Prospective Study (MCPS): 136,725 participants
- Mayo Clinic Project Generation cohort (MAYO-RGC): 80,941 participants
- BangladEsh Longitudinal Investigation of Emerging Vascular and nonvascular Events (BELIEVE): 69,663 participants
- Colorado Center for Personalized Medicine biobank (CCPM): 56,173 participants
- UCLA ATLAS Precision Health Biobank (ATLAS): 45,533 participants
- Mount Sinai BioMe BioBank (BioMe): 45,440 participants
- University of Pennsylvania Penn Medicine BioBank (PMBB): 25,785 participants
- Dallas Biobank Study (DBS): 13,979 participants
- Malmö Diet and Cancer Study (MDCS): 5,146 participants
Analyses of additional cardiometabolic phenotypes also included 4,755 participants from the Indiana University School of Medicine study (Indiana-CLDB) and 53,325 participants from the South Asia Biobank.
The MCPS included residents of the Coyoacán and Iztapalapa districts of Mexico City. This cohort represents a genetically admixed population with Indigenous American, European and African ancestry components.
All participating studies received approval from the appropriate ethics committees, and all participants provided informed consent. Ethical approvals included the following:
- UK Biobank: North West Centre for Research Ethics Committee, approval 11/NW/0382; UK Biobank application 26041.
- GHS: Geisinger institutional review board approval 2006-0258.
- MCPS: Approvals from the Mexican National Council of Science and Technology, Mexican Ministry of Health and Central Oxford Research Ethics Committee, including C99.260.
- MAYO-RGC: Mayo Clinic IRB protocol 09-007763.
- BELIEVE: Approvals from the Bangladesh Medical Research Council, National Heart Foundation Hospital and Research Institute, icddr,b and Bangabandhu Sheikh Mujib Medical University.
- CCPM: University of Colorado Anschutz Medical Campus IRB protocol 15-0461.
- ATLAS: UCLA IRB approval 17-001013.
- BioMe: Icahn School of Medicine at Mount Sinai PPHS IRB approvals 11-01139 and 21-01743.
- PMBB: University of Pennsylvania IRB protocol 813913.
- DBS: UT Southwestern IRB study STU 022011-116.
- MDCS: Lund University Ethics Committee approval LU 51-90 and Regional Board of Ethics in Lund approval Dnr 2016/479.
- Indiana-CLDB: Indiana University IRB protocol 1105005445.
- South Asia Biobank: Institutional Research Ethics Committee approval 18IC4698.
Clinical Phenotypes and Biomarkers
Blood Lipids and Cardiometabolic Measurements
Blood lipid levels, glycaemic markers, transaminases, C-reactive protein and blood pressure were obtained through standardized clinical assessments or electronic health record (EHR) extraction, depending on the cohort.
In population-based cohorts, measurements were collected during standardized visits at assessment centres. In health-system cohorts, measurements were retrospectively extracted from EHRs. When multiple clinical measurements were available, participant-specific median values were used.
Most cohorts measured blood lipids using routine clinical chemistry methods. In the MCPS and BELIEVE cohorts, lipid levels were derived from nuclear magnetic resonance metabolomics using the Nightingale Health platform.
Lipid measurements were adjusted for medication use by dividing values by class-specific correction factors derived from previously reported clinical trials.
The TG:HDL ratio was calculated using triglyceride and HDL cholesterol values obtained from the same blood draw in five epidemiological cohorts, representing 635,115 participants. In six EHR-based cohorts, representing 397,001 participants, median triglyceride and HDL cholesterol values were calculated separately across clinical encounters, and the ratio of these median values was used.
Using median values from multiple clinical encounters is a standard approach in EHR-based genetic research. This method captures a participant’s typical lipid profile while avoiding exclusion of individuals whose valid triglyceride and HDL measurements were not collected concurrently.
Sensitivity analyses were performed using only TG:HDL ratios calculated from triglyceride and HDL measurements obtained during the same blood draw. In the UK Biobank, analyses were also stratified by fasting status. Fasting samples were defined as those collected at least 6 hours after the last meal, whereas non-fasting samples were collected no more than 2 hours after the last meal.
For participants receiving antihypertensive medication, 15 mmHg was added to systolic blood pressure and 10 mmHg was added to diastolic blood pressure.
Liver Biopsy and Histopathology
The GHS bariatric surgery cohort included 3,599 participants of European ancestry who underwent bariatric surgery. Intraoperative wedge liver biopsies were collected approximately 10 cm to the left of the falciform ligament before manipulation of the liver or stomach.
For clinical histopathology, tissue was fixed in 10% neutral buffered formalin and stained with haematoxylin and eosin for general assessment and Masson’s trichrome for fibrosis evaluation. Remaining tissue was preserved in RNAlater or snap-frozen in liquid nitrogen for research analyses.
Two pathologists independently reviewed the biopsies. Histological scoring followed the Nonalcoholic Steatohepatitis Clinical Research Network system:
- Steatosis: grade 0, less than 5%; grade 1, 5% to less than 34%; grade 2, 34% to less than 67%; and grade 3, more than 67% parenchymal involvement.
- Lobular inflammation: grade 0, none; grade 1, mild; grade 2, moderate; and grade 3, severe.
- Hepatocyte ballooning: grade 0, none; grade 1, few ballooned cells; and grade 2, many or prominent ballooned cells.
- Fibrosis: staged from 0, no fibrosis, to 4, cirrhosis.
The metabolic dysfunction-associated steatotic liver disease (MASLD) activity score was calculated by adding the steatosis, lobular inflammation and hepatocyte ballooning scores.
Imaging Phenotypes
MRI Scans
MRI data were obtained from the UK Biobank imaging cohort. Liver fat was quantified using two-dimensional abdominal MRI scans. Some participants were scanned using a Dixon gradient-echo protocol, while scans obtained from 2016 onward used the IDEAL protocol, which applies iterative decomposition of water and fat with echo asymmetry and least-squares estimation.
The images had an in-plane pixel size of 2.5 × 2.5 mm and a slice thickness of 6 mm.
Whole-body MRI scans extended from the neck to the knees and used a two-point Dixon spoiled gradient-echo T1-weighted protocol. Echo times were 2.39 ms out of phase and 4.77 ms in phase, with a repetition time of 6.69 ms and a 10-degree excitation flip angle.
Image acquisition was divided into six anatomical stages, each representing a separate three-dimensional volume. In-plane resolution was 2.23 × 2.23 mm, while out-of-plane resolution ranged from 3.00 to 4.50 mm. All MRI scans were acquired using Siemens MAGNETOM clinical scanners.
MRI-Derived Body Fat and Liver Fat Measures
An automated segmentation model was used to quantify major adipose tissue depots from Dixon MRI data. Fat was divided into visceral and subcutaneous compartments. Visceral fat was further classified as mediastinal or abdominal, while subcutaneous fat was divided into upper-body, abdominal and gluteofemoral compartments.
Fat volume was calculated by summing voxel-level fat fractions within each segmented region. The visceral-to-gluteofemoral fat ratio was calculated as abdominal visceral fat volume divided by subcutaneous gluteofemoral fat volume. The abdominal-to-gluteofemoral subcutaneous fat ratio was calculated using the corresponding total volumes.
All image processing and phenotype extraction were performed on the UK Biobank Research Analysis Platform.
Liver fat was measured as the proton density fat fraction (PDFF), representing the proportion of liver signal attributable to fat. A deterministic image-processing algorithm segmented the liver and used multithresholding to exclude blood vessels. PDFF was calculated as the fat signal divided by the combined fat and water signal.
To account for systematic differences between gradient-echo and IDEAL measurements, mean values were aligned across modalities. Imaging modality was also included as a covariate in genetic analyses.
Bioelectrical Impedance Analysis and DEXA
Baseline body composition was measured using the Tanita BC 418ma Body Fat Analyzer. Fat mass and fat-free mass were obtained from UK Biobank fields 23100 and 23101. Fat-free mass was defined as total body weight minus fat mass. Body fat percentage was calculated using fat mass and baseline weight.
DEXA measurements were obtained using the GE-Lunar iDXA instrument. Lean mass and android-to-gynoid fat mass measurements were obtained from UK Biobank fields 23280, 23245 and 23262.
Disease Definitions
Type 2 Diabetes
Type 2 diabetes cases were identified using established clinical and research algorithms. A participant was classified as a case if at least one of the following criteria was met:
- At least one inpatient or two outpatient EHR entries with type 2 diabetes codes, including ICD-10 codes E11 or O24.1, equivalent ICD-9 codes or diabetes recorded as a cause of death.
- Glycaemic biomarkers, including HbA1c, random glucose or fasting glucose, within the diabetic range.
- Algorithmically defined type 2 diabetes based on self-reported medical history and medication use.
Participants meeting criteria for type 1 diabetes, including ICD-10 codes E10 or O24.0 or algorithmic classification based on self-reported information and medication use, were excluded from the case group.
Controls did not meet the type 2 diabetes criteria and were excluded if they had EHR evidence of any diabetes type, a family history of diabetes or glycaemic biomarkers in the prediabetic range.
Liver Disease
Liver disease cases, including MASLD and cirrhosis, were identified using established criteria. Cases required at least one of the following:
- Liver disease documented during at least one inpatient encounter, two or more outpatient encounters or as a cause of death.
- A self-reported liver disease diagnosis at baseline.
- A documented history of surgical or medical treatment for liver disease.
Controls were excluded if they had another liver disease, an outpatient encounter for the liver disease under investigation, elevated ALT levels above 25 IU l−1 in women or 33 IU l−1 in men, or ascites attributed to liver disease.
Coronary Artery Disease
Coronary artery disease (CAD) cases were identified using EHR diagnoses, self-reported medical history and procedure records. Participants were classified as cases if they had:
- CAD or myocardial infarction documented in at least one inpatient encounter, two or more outpatient visits or as a cause of death.
- A self-reported diagnosis of CAD or myocardial infarction at baseline.
- A history of CAD-related procedures, including coronary artery bypass grafting or percutaneous coronary intervention.
Controls were excluded if they had a family history of CAD based on EHR or self-reported information.
Epidemiological Analysis
Associations between the TG:HDL ratio and clinical, metabolic, imaging and histopathological phenotypes were assessed using correlation and linear regression models. TG:HDL ratio values were categorized into percentiles.
Cox proportional hazards models were used to evaluate the relationship between the TG:HDL ratio and incident type 2 diabetes, myocardial infarction, MASLD and liver cirrhosis. Analyses were restricted to individuals who did not have the relevant outcome at or before baseline.
DNA Preparation and Exome Sequencing
Genomic DNA libraries were created by enzymatically fragmenting high-molecular-weight DNA to an average size of approximately 200 bp. Unique 10-bp asymmetric barcodes were added during amplification to enable multiplexed exome capture and sequencing.
Barcoded libraries were pooled in equimolar quantities and captured using either a modified xGen Exome Research Panel or a modified Human Comprehensive Exome Panel. Enriched libraries were amplified, quantified and sequenced on Illumina HiSeq 2500 or NovaSeq 6000 instruments using 75-bp paired-end reads.
Read Mapping and Variant Calling
FASTQ files were generated from Illumina image data using bcl2fastq v2.19.0. Reads were aligned to the GRCh38 reference genome with BWA-MEM v0.7.15 using an alt-aware configuration. Duplicate reads were marked, and per-read annotations were added.
Variants were identified using a Parabricks-accelerated implementation of DeepVariant v0.10 with a custom model for single-nucleotide variants and short insertions and deletions. Individual VCF files were jointly genotyped with GLnexus v1.4.3. PLINK v1.9 was used to convert the resulting multisample VCF files for downstream analyses.
Exome Sequencing Quality Control
Quality-control criteria included sample missingness below 10%, variant missingness below 10%, a ±10-bp buffer around target regions, a low-heterozygosity Hardy–Weinberg equilibrium threshold of P > 1 × 10−100 and an excess-heterozygosity threshold of P > 1 × 10−30.
Variant Annotation
Variants were annotated using the Variant Effect Predictor v100.4 with Ensembl release 100 protein-coding transcript models. A single functional consequence was assigned using the canonical transcript, selected with MANE, APPRIS and Ensembl canonical transcript annotations.
Nonsense-mediated decay was predicted with the VEP NMD plugin. Variants were considered likely to escape nonsense-mediated decay when they occurred in the final exon, within 50 bases upstream of the penultimate exon, within the first 100 coding bases or in an intronless transcript.
Genetic Association Analyses
Genetic association analyses were conducted separately within each cohort using the REGENIE whole-genome regression framework. Analyses included:
- A genome-wide association study of common TOPMed-imputed variants with an alternate allele frequency of at least 1%.
- Gene-based association testing of rare coding variants.
Directly genotyped variants used in the first REGENIE step had an alternate allele frequency of at least 1%, missingness below 10%, Hardy–Weinberg equilibrium P > 10−15 and passed linkage disequilibrium pruning.
Regression models included age, age2, sex, age-by-sex and age2-by-sex interactions, the first 10 principal components from common variants, the first 20 principal components from rare variants and cohort-specific sequencing batch covariates. A genome-wide leave-one-chromosome-out polygenic score was also included to account for relatedness and residual population stratification.
The primary analysis pooled participants across ancestries. Sensitivity analyses were stratified by European, admixed American, South Asian, African and East Asian ancestry.
For non-pseudoautosomal regions of chromosome X, dosage compensation was modeled by coding homozygous-reference males as 0, hemizygous males as 2 and heterozygous males as missing.
To distinguish rare-variant signals from common-variant effects in the same chromosome, gene-burden analyses were additionally adjusted for common variants identified through fine-mapping. Ancestry-specific fine-mapping was also used in a sensitivity analysis.
For 59 independent gene-burden signals, additional analyses were stratified by ancestry, fasting status and sex. Leave-one-covariate-out analyses demonstrated correlations of r > 0.99 for triglycerides, HDL cholesterol and the TG:HDL ratio, indicating minimal influence from individual age and sex covariates.
Common-Variant Analysis and Fine-Mapping
A GWAS of common variants with a minor allele frequency of at least 1% was performed for the TG:HDL ratio. Cohort-specific results were combined using fixed-effect inverse-variance-weighted meta-analysis.
Variants reaching genome-wide significance, defined as P < 5 × 10−8, were fine-mapped using SuSiE v0.12.35. Candidate regions were defined around variants with P < 1 × 10−6, separated by at least 100 kb and extended to nearby recombination hotspots. Regions were expanded when necessary and merged when their 100-kb buffers overlapped.
An exact pooled linkage disequilibrium matrix was calculated using participants included in the association analyses. Cohort-specific covariance matrices were weighted by their degrees of freedom and combined across cohorts.
SuSiE generated credible sets containing variants with a high probability of being causal. Posterior inclusion probabilities were assigned to each variant, and the smallest set accounting for 95% of the cumulative probability was defined as the 95% credible set. The variant with the highest posterior inclusion probability at each locus was designated the sentinel variant.
Rare-Variant Gene-Burden Analysis
Variants predicted to cause frameshifts, premature stop codons or disruption of canonical splice donor or acceptor sites were classified as predicted loss-of-function variants.
Missense variants were evaluated using ESM1vp, SIFT, PolyPhen2-HumDiv, PolyPhen2-HumVar, MutationTaster and LRT. Gene-burden tests combined variants according to allele-frequency thresholds and predicted functional impact.
Four maximum alternate allele-frequency thresholds were tested: singleton, 0.0001, 0.001 and 0.01. Six variant groupings were evaluated:
- Predicted loss-of-function variants.
- Predicted loss-of-function plus deleterious missense variants identified by at least one of five non-ESM1vp algorithms.
- Predicted loss-of-function plus missense variants predicted deleterious by all five non-ESM1vp algorithms.
- Predicted loss-of-function plus missense variants predicted deleterious by ESM1vp.
- Predicted loss-of-function plus missense variants predicted deleterious by all five non-ESM1vp algorithms and ESM1vp.
- Predicted loss-of-function plus all missense variants.
Each gene generated 24 burden exposures. Their P values were combined using BURDEN-ACAT. The exome-wide significance threshold was P < 1.04 × 10−7, based on Bonferroni correction for approximately 20,000 genes and 24 burden tests per gene.
Within-chromosome conditional analyses were performed by adjusting each gene-burden signal for other burden signals on the same chromosome.
Tissue-Enrichment Analysis
Gene expression data from the GTEx version 8 data freeze were used to identify tissues enriched for expression of genes associated with the TG:HDL ratio.
Median expression levels were calculated for each gene and tissue after RNA-sequencing quality control. Expression values were standardized within tissues and then across tissues for each gene. Expression-enhanced tissues were defined as those with expression at least 6 standard deviations above the gene-specific median across tissues.
For gene-burden associations, burden-test P values were aggregated by gene using the Cauchy combination test. Tissue enrichment odds ratios were estimated using Firth-corrected logistic regression, excluding primary sexual and reproductive tissues.
For common-variant analyses, tissue enrichment was assessed using protein-coding genes prioritized through variant-to-gene analyses. Gene length was included as a covariate to reduce potential confounding.
Liver RNA Sequencing and eQTL Analysis
Liver RNA sequencing was performed on samples from 1,946 GHS participants who underwent perioperative wedge liver biopsy during bariatric surgery.
RNA Preparation and Sequencing
Total RNA was extracted using the NEBNext Poly(A) mRNA Magnetic Isolation Module and NEBNext Ultra II Directional RNA Library Prep Kit. Libraries were amplified using Kapa HiFi polymerase and custom barcoded primers, then sequenced on the Illumina NovaSeq 6000 platform with paired-end 75-bp reads.
Mean sequencing depth was 72 million reads per sample, with a median of 68 million reads. Ninety-three percent of samples produced at least 50 million reads, and 99% produced more than 45 million reads.
RNA-Sequencing Quality Control
Reads were aligned to the GRCh38/hg38 reference genome using STAR v2.5.3a. Optical duplicates were marked with Picard. Gene-level expression was quantified using RNA-SeQC and GENCODE release 32 annotations.
Only uniquely mapped, properly paired reads with alignment distances of no more than six and reads fully contained within exon boundaries were retained. Expression values were normalized using the trimmed mean of M values method. Genes were retained if they reached at least 0.1 TPM and six unnormalized reads in at least 20% of samples.
eQTL Analysis and Colocalization
eQTL analyses followed the GTEx workflow. Models included age, sex, four common-variant principal components and 100 gene-expression principal components to account for potential batch effects. Gene expression values were transformed using a rank-inverse normal transformation.
Fine-mapped TG:HDL ratio sentinel variants were integrated with liver eQTL data. Variants in strong linkage disequilibrium with sentinel variants, defined as r2 > 0.8, were also included. Bayesian colocalization was performed within 500-kb windows using coloc v5.2.3 to estimate the probability that liver eQTL and TG:HDL ratio associations shared a causal variant.
Proteomics and Protein QTL Analysis
UK Biobank plasma proteomics data were obtained from the UK Biobank Plasma Proteomics Project. In the GHS cohort, serum samples from 9,941 participants were analyzed using the Olink Explore 3072 assay.
Proteins were measured across cardiometabolic, inflammation, neurology and oncology panels. Samples underwent principal-component outlier removal, sample-swap exclusion and NPX intensity normalization.
Protein levels were rank-inverse normal transformed and adjusted for age, sex, age2, age-by-sex and age2-by-sex interactions. Final pQTL results were generated by inverse-variance-weighted meta-analysis of UK Biobank and GHS data.
Variant-to-Gene Prioritization and Pathway Analysis
Genes were prioritized from the TG:HDL ratio GWAS using three criteria:
- Colocalization with fine-mapped GTEx eQTL peaks in any tissue, with r2 ≥ 0.8.
- Colocalization with fine-mapped pQTL peaks, with r2 ≥ 0.8.
- Physical proximity to the nearest gene.
Pathway enrichment was assessed using DEPICT gene-level pathway z-scores. The 59 target genes were mapped to Ensembl identifiers and intersected with the DEPICT gene set. Stouffer scores were used to calculate pathway enrichment, with Bonferroni-corrected significance defined as P < 5.9 × 10−6.
Mouse Models and Experimental Procedures
C57BL/6NTac mice and Cas9-Ready mice expressing Cas9 under the CAG promoter were used. Male mice were housed under a 12-hour light–dark cycle at 22 ± 1 °C or under thermoneutral conditions at 30 °C. Animals had free access to food and water and were fed either standard chow or a high-fat, high-fructose diet.
Mice were 6 to 11 weeks old at study initiation. Age- and strain-matched littermates were used whenever possible. Experimental groups were balanced by baseline body weight, body composition and lipid levels. Investigators were blinded during serum chemistry, liver lipid, protein and DNA-editing analyses, although blinding was not possible during live treatment procedures.
Body weight was measured weekly, and body composition was assessed in awake mice using EchoMRI. Plasma was collected at defined time points after adeno-associated virus administration. Only male mice were studied, which may limit the generalizability of the animal findings to females. All animal procedures were approved by the Regeneron Pharmaceuticals Institutional Animal Care and Use Committee.
AAV Administration
AAV8 vectors encoding short hairpin RNAs were used for liver-targeted knockdown. Vectors were diluted in saline and administered intravenously at 2.5 × 1011 viral genomes per mouse.
An Hpn-targeting AAV encoded a 1:1 mixture of two shRNAs. A non-targeting scramble shRNA served as the control, and saline was used as a vehicle control.
For CRISPR-based studies, AAV8 vectors carrying multiple guide RNAs targeting Flcn, Fnip1, Fnip2 or combined Fnip1 and Fnip2 were administered intravenously at 2.5 × 1010 viral genomes per mouse. Each construct contained five guide RNAs driven by individual U6 promoters.
siRNA Administration
GalNAc-conjugated siRNAs targeting Flcn or a non-targeting control were administered subcutaneously at 10 mg kg−1 every 10 days.
DNA Editing and Amplicon Sequencing
Genomic DNA was extracted from liver tissue. Target-specific oligonucleotides were designed to generate amplicons no longer than 350 bp. Barcoded adapters were added, and PCR products were pooled and purified using AMPure XP reagent.
Libraries were sequenced on the Illumina MiSeq platform using a 2 × 300 read kit. Paired reads were merged with PEAR and aligned to the mouse mm10 genome using Bowtie2. Each sample received at least 20,000 merged reads across the expected guide-cleavage site.
Insertions, deletions and base changes within 20 bases upstream or downstream of the expected cleavage site were classified as CRISPR-induced modifications. Editing efficiency was calculated by comparing modified reads with wild-type reads.
Liver Lipid Analysis and Glycaemic Testing
Blood was collected into EDTA tubes and plasma was isolated by centrifugation. Total cholesterol, LDL cholesterol, HDL cholesterol, non-esterified fatty acids, albumin, total protein and ALT were measured using the ADVIA Chemistry XPT analyzer. Non-HDL cholesterol was calculated by subtracting HDL cholesterol from total cholesterol.
For liver lipid analysis, snap-frozen liver samples were weighed, homogenized in chloroform:methanol and separated into organic and aqueous phases. Extracted lipids were solubilized and triglyceride and cholesterol concentrations were measured enzymatically and normalized to wet tissue weight.
Insulin tolerance tests were performed after a 4-hour fast. Insulin was administered intraperitoneally at 0.75 U kg−1. Blood glucose was measured at baseline and 30, 60, 90 and 120 minutes. Plasma insulin was measured using a mouse insulin ELISA.
Gene Expression and RNA Sequencing in Mice
RNA was extracted from tissue samples preserved in RNAlater. Genomic DNA was removed, and RNA was reverse transcribed into complementary DNA. Quantitative PCR was performed using TaqMan assays, with relative expression calculated using the 2−ΔΔCt method.
For mouse liver RNA sequencing, samples included chow-fed mice, high-fat, high-fructose diet vehicle controls and mice treated with Flcn-targeting AAV8 guide RNAs. RNA-sequencing libraries were generated using the KAPA mRNA HyperPrep Kit and sequenced on the Illumina NovaSeq 6000 platform.
Reads were mapped to the mouse GRCm38 reference genome. Gene-level counts were analyzed with DESeq2. Benjamini–Hochberg correction was applied, and significant genes were defined using a false discovery rate below 0.05 and a fold change greater than 2.
Single-Cell RNA Sequencing
Five human liver single-cell RNA-sequencing datasets were obtained from the Gene Expression Omnibus: GSE136103, GSE168933, GSE115469, GSE158723 and GSE185477.
Doublets were identified and removed using Scrublet. Cells with fewer than 200 detected genes or mitochondrial counts exceeding 10% were excluded. Datasets were merged, normalized to 10,000 counts per cell and log-transformed using Scanpy.
The 5,000 most variable genes were identified, and principal-component analysis was performed. Batch effects were corrected using Harmony. A shared nearest-neighbour graph, UMAP embedding and Leiden clustering were generated. Cell types were annotated using canonical markers, with myeloid and lymphoid compartments further subclustered.
Mouse liver single-cell data were obtained from study GSE156052. Author-provided UMAP coordinates and cell-type annotations were used without modification.
Protein Detection by Liquid Chromatography–Mass Spectrometry
Liver protein lysates were prepared in SDS buffer containing protease inhibitors. Proteins were reduced, alkylated and digested into peptides using trypsin and Lys-C.
Stable isotope-labelled peptides for FLCN and FNIP2 were added for validation. Peptides were separated on a C18 column and analyzed using parallel reaction monitoring on a Thermo Scientific Exploris 480 mass spectrometer.
Target peptides were quantified from extracted ion chromatograms. Measurements were normalized to histone H4C1 to account for differences in sample input.
FNIP1 Knockdown in Primary Human Hepatocytes
Cryopreserved primary human hepatocytes were obtained from Thermo Fisher Scientific and plated on collagen-coated 12-well plates. After cells formed a monolayer, they were maintained in serum-free Williams’ E medium.
Cells were transfected with 100 nM FNIP1-targeting or non-targeting siRNAs using Lipofectamine RNAiMAX. Cells were collected after 96 hours for RNA extraction, gene-expression analysis and immunoblotting.
Immunoblotting
Cell lysates were prepared in RIPA buffer with protease inhibitors. Equal protein quantities were separated by SDS–PAGE and transferred to PVDF membranes. Membranes were probed with antibodies against human FNIP1 and HSP90, and signals were detected using enhanced chemiluminescence.
Statistical Analysis and Reproducibility
Cellular and animal data were analyzed using Microsoft Excel and GraphPad Prism 10. Data are presented as mean ± standard error of the mean or mean ± 95% confidence intervals.
Group means were compared using one-way or two-way analysis of variance. A two-sided P value below 0.05 was considered statistically significant.
Experiments involving AAV8 guide RNAs targeting Flcn, Fnip1 and Fnip2 in high-fat diet-fed mice for up to 13 weeks, primary human hepatocyte Fnip1 knockdown and AAV-shRNA targeting Hpn were independently performed twice and produced similar results.
Long-term 30-week inhibition studies, chow-diet experiments, thermoneutrality experiments and GalNAc-conjugated siRNA studies were performed once. Their results were consistent with findings from the AAV8 guide-RNA experiments.
Additional information on research design is available in the Nature Portfolio Reporting Summary associated with the study.
Source: www.nature.com


