TRACE-CpG Logo TRACE-CpG
Trans-tissue Regression Approximation of Central Epigenome from Correlated peripheral & Gray-matter

About TRACE-CpG

DNA methylation (DNAm) is highly tissue-specific, and in neuropsychiatric epigenetics, DNAm profiles measured in the human brain are therefore the most direct substrate of interest. However, because access to brain tissue from living individuals is rare, many studies rely on peripheral tissues such as blood, buccal epithelial, and saliva as proxy measures. IMAGE-CpG2 and COMPASS address a shared central question—which peripheral DNAm signals reflect DNAm variation in the brain—by providing empirical brain–periphery correlation maps at single-CpG and gene/locus levels. TRACE-CpG approaches the same challenge from a complementary perspective. Rather than querying correlation alone, it predicts brain DNAm by tracing brain DNAm states from peripheral DNAm at CpG sites where peripheral measurements capture sufficient information about brain variation.

TRACE-CpG is a linear mixed-effects framework developed from the paired brain–periphery EPIC-array data assembled in IMAGE-CpG2 (DB6–DB8). Building on the brain–periphery correlations characterized through IMAGE-CpG2 and COMPASS, it models the relationship between brain and peripheral DNAm at each matched CpG site. In the web tool, users upload their own peripheral DNAm data as a matrix, together with the age and sex metadata required for prediction. The tool then applies pre-estimated, fixed model coefficients to output predicted pan-brain DNAm values. Predictions are provided only for brain epigenome surrogate sites (BESS)—CpG sites that show sufficient brain–periphery correlation and adequate predictive performance in internal evaluation (marginal R² > 0.2)—covering approximately 150,000 sites for blood and buccal epithelial cells and approximately 60,000 sites for saliva.

By extending cross-tissue correlation into predicted brain DNAm values, TRACE-CpG supports brain-informed interpretation of peripheral epigenetic findings. These predictions are most useful for prioritizing and interpreting individual loci with strong brain relevance, and remain limited to CpG sites with sufficient cross-tissue signal.

Instructions & Notes

Download example input files
Includes GSE214901_metadata.csv and GSE214901_beta.csv for testing and formatting validation.

Input requirements

  • Array type: Illumina Infinium MethylationEPIC v1 or EPIC v2 (IDAT-derived).
    Not currently supported: HumanMethylation450 (450K). The 450K array uses a different probe set and manifest/annotation structure than EPIC platforms, and direct compatibility may be limited.
  • Input format: a beta-value matrix with values in [0, 1] (rows = CpGs, columns = samples), plus a metadata CSV (Sample_ID, Age, Sex).
  • Preprocessing: Data should be background-corrected and normalized from IDAT files.
  • Probe/sample QC: Remove poor-quality samples/probes (e.g., low detection p-values / low beadcount, extreme outliers) prior to upload.
  • (Equivalent preprocessing pipelines are acceptable as long as they produce normalized EPIC beta values.)

Recommended preprocessing pipelines
The following pipeline is recommended by the Psychiatric Genomics Consortium (PGC) PTSD EWAS working group and takes data from IDAT files through normalization.

https://github.com/skatrinli/PGC-PTSD-EWAS-EPIC_QC
Scripts: PGCpipeline-1-QualityChecks.R, PGCpipeline-2-Preprocessing-Normalization.R

This pipeline is used for both EPIC v1 and EPIC v2. For EPIC v2, apply the same normalization, and additionally handle the collapsing of replicate (suffix) probes to their v1-equivalent CpG identifiers, referring to https://github.com/skatrinli/EPICv2_pipeline for the suffix-merging step.

Notes
If you use minfi (or similar), please ensure you apply an appropriate normalization method such as ssNoob and export beta values.


TRACE-CpG output — important notes

  • TRACE-CpG reports predictions only for CpG sites that are shared between EPIC v1 and EPIC v2. Array-specific CpGs (unique to EPIC v1 or EPIC v2) are not included.
  • Output is limited to CpG sites that pass our cross-tissue and prediction-performance filters (marginal R² > 0.2 and FDR < 0.05 from permutation testing); therefore, not all input CpGs will be returned.
  • Predictions represent "pan-brain" DNAm estimates (not region-specific brain DNAm).
  • All values are reported as beta values (β; 0–1). If needed, users can convert β to M values for downstream analyses.
  • Predicted values are statistical estimates, not direct measurements of brain DNAm.
  • Runtime depends on sample size; typical runs finish in 8-10 seconds for ~35 samples.
  • Important: Age and sex are used as predictors in TRACE-CpG. Therefore, predicted brain DNAm values partially reflect these inputs. Association tests of the predicted values with age and/or sex may be inflated (circular analysis) and should be avoided. In downstream case–control analyses, adjusting for age and sex as covariates may still be appropriate to reduce confounding, depending on the study design and research question.

Select peripheral tissue for prediction:

Sample metadata (CSV)

Required columns: Sample_ID, Age, Sex (M/F), Tissue (optional)

DNAm beta values (CSV)

First column: CpG_ID; others: sample beta values (0–1).

Please upload Phenotype file first


Reference

Nishitani S, Shinozaki G et al., (2026). XXXX, bioRxiv

Contact

Shota Nishitani, PhD — Basic Life Research Scientist, Department of Psychiatry and Behavioral Sciences, Stanford University School of Medicine (ORCID: 0000-0002-6234-295X)
Gen Shinozaki, MD — Associate Professor, Department of Psychiatry and Behavioral Sciences, Stanford University School of Medicine (ORCID: 0000-0001-6129-2789)

© AsahiKato, Shinozaki Lab, Stanford University School of Medicine