Whole genome
Broadest view of inherited and tumor-acquired change.
- SNVs and indels
- Copy number and structural variation
- Non-coding and breakpoint context
- Best paired with matched normal
Every contribution should enter the Foundation's data as a complete, longitudinal, analysis-ready patient record, not as another folder of disconnected files.
This page is a design. It describes how the Foundation intends to collect, clean, govern and share the data it will hold. None of it is running today, and the legal statements in it are drafts for counsel and IRB review.
The unit of analysis is not a sequence file. It is a patient, a disease state, a lesion, a specimen, a treatment exposure, and an outcome, linked over time.
Diagnosis, disease burden, performance, comorbidities, every treatment, dose, hold, surgery, radiation, response and toxicity.
Lesion identity, anatomic site, collection time, preservation, tumor content, necrosis, cold ischemia and pathology review.
Instrument-level files whenever available, retained immutably so future methods can reprocess the original evidence.
Original reports, spreadsheets, calls and interpretations are preserved beside the harmonized analysis, not substituted for it.
Early molecular response, imaging, symptoms, quality of life, progression, resistance and subsequent interventions.
Who generated each file, with which instrument, reference, software, parameters, annotation set and quality status.
The record should preserve the strengths and limitations of each assay, then connect them into a coherent view rather than flattening them into one vendor report.
Broadest view of inherited and tumor-acquired change.
Efficient depth across protein-coding genes and a strong cross-validation layer.
Useful clinical snapshots with fast interpretation, but limited genomic breadth.
Shows which genes and pathways are active at a particular disease state.
Resolves transcript structure that short reads can fragment or misassign.
Separates tumor, immune and stromal states that bulk measurements average together.
One timepoint is a snapshot. Repeated sequencing turns the disease into a movie.
Serial tissue and blood allow within-patient comparisons across diagnosis, treatment, recurrence and metastasis. Multiple lesions collected at the same procedure reveal spatial heterogeneity as well as temporal evolution. The timepoints below are from the illustrative catalog.
Baseline biology before systemic therapy.
Residual disease after first treatment pressure.
Biology after completion of standard therapy.
Fresh-frozen and FFPE material for deeper analysis.
Independent nodules with shared matched normal.
Repeat the core assays and add high-value modalities where tissue permits.
*An expected WGS component in the illustrative catalog was not present in the delivered folder. The record would mark it expected-but-missing, not as a negative result.
Stable identifiers and validated relationships prevent the most damaging class of errors: mispaired normals, merged lesions, ambiguous timepoints, lost aliquot history and results detached from treatment context.
The graph keeps biological replicates, technical replicates, multiple nodules, shared normals and repeated assays distinct while making their relationships directly queryable.
A contribution would count as complete only after it passes a data contract and produces versioned, quality-scored, analysis-ready outputs.
Inventory every expected and received file; verify checksums, byte size, compression and readability.
Fingerprint samples, confirm patient identity, tumor-normal pairing, lesion IDs, sex and contamination.
Map site, timepoint, preservation, assay, platform, panel, read group and treatment context to controlled terms.
Coverage, duplication, insert size, mapping, RNA quality, tumor purity, read structure and modality-specific checks.
Reprocess DNA and RNA against a versioned common reference and annotation bundle while retaining source alignments.
Generate SNV/indel, CNV, SV, fusion, expression, isoform and other derived outputs with reproducible workflows.
Normalize coordinates and identifiers; add transcript, gene, protein, pathway and clinical annotations with versions.
Measure concordance, retain disagreement, identify report-only evidence and never convert missing into negative.
Apply de-identification, access tier, consent and data-use tags. Preserve intervals in longitudinal research views. Section 06 describes the machine-readable stewardship layer.
Release immutable raw, harmonized and analysis-ready layers through versioned snapshots.
| Data type | Work completed at contribution | Research-ready result |
|---|---|---|
| DNA | Common reference; read-group validation; tumor-normal identity; SNV/indel, copy number and structural-variant processing; variant normalization. | Comparable variant and copy-number tables across laboratories, panels and timepoints, plus source-native calls. |
| Short-read RNA | Common genome and transcript annotation; gene-level counts; expression normalization; fusion analysis; quality and batch covariates. | Counts, TPM, fusions, pathway-ready matrices and explicit assay/batch metadata. |
| Long-read RNA | HiFi QC; transcript alignment; isoform collapse and classification; junction and fusion reconciliation with short reads. | Full-length isoform and transcript-event tables linked to expression and genomic breakpoints. |
| Single-cell / spatial | Barcode and UMI QC; doublet and ambient-RNA assessment; raw-count preservation; cell and region annotations; batch flags. | Portable matrices with curated cell labels, embeddings, QC fields and spatial coordinates where available. |
| Protein / peptide | Standard protein and phosphosite IDs; FDR and intensity normalization; peptide-to-protein and variant-peptide mapping; raw-spectra status. | Comparable protein, phosphosite and immunopeptide features with confidence and provenance. |
| Clinical / imaging | Medication, lab, disease, toxicity and response vocabularies; lesion IDs; DICOM de-identification; treatment intervals and outcomes. | Queryable treatment-exposure, lesion-response, toxicity and outcome tables synchronized to molecular timepoints. |
The design is one portable consent and stewardship layer for trials, institutions and patient-mediated contribution. Every future trial would contribute through that same layer. Each patient would choose how their data may be used, and every permission would become a machine-readable rule that follows the longitudinal patient graph.
Concept draft. The legal and regulatory statements in this section are for counsel and IRB review. Portability depends on IRB approval, reliance or contribution agreements, and the privacy law of each jurisdiction. The full legal basis table is in the working specification, which is not published here.
Trial-mediated
Institution-mediated
Patient-mediated
Clinical findings would be returned only through an approved clinical process. Nothing on this page is medical advice.
Existing cohorts
Open question. Federally funded genomic data must still be deposited in a designated federal repository under the applicable genomic data sharing policy, and the Foundation cannot replace those obligations. The design needs an approved-repository or trusted-partner arrangement so those datasets can be reached through the same request.
Multimodal models for rare pediatric cancers cannot be judged on data they were trained on. The design keeps a sequestered evaluation set for exactly that purpose, and patients and contributors would see it as a named use rather than a side effect.
A parent or legally authorized representative gives permission. The child gives assent when developmentally appropriate. Re-consent would be requested when the participant reaches the legal age of adulthood. The patient would be able to review scopes, change future permissions, see access history and request withdrawal at any time.
Withdrawal is prospective. The Foundation would stop new use and future release when the withdrawal takes effect, subject to legal and research-record duties. Data already released publicly, included in completed analyses or published cannot be recalled. If a participant dies, the choices on record continue to govern the data, and a parent or legal representative may still request prospective withdrawal.
The full legal basis table, with counsel flags, is in the working specification and is not published here.
Technical and governance precedents: Count Me In; NCI Genomic Data Commons and dbGaP controlled access; GA4GH Data Use Ontology and Data Passports; Sage Bionetworks Elements of Informed Consent; ICGC ARGO consent framework; UK Biobank governance; and Break Through Cancer Data Science Hub. These are precedents, not partners, and not endorsements of this proposed design.
The illustrative catalog includes raw reads, aligned files, reports, spreadsheets, multiple laboratories, FFPE and flash-frozen tissue, shared normals, tumor-only data, report-only deliveries and expected files that have not yet arrived.
"No prep work" should mean no foundational wrangling. Researchers still choose the scientific question, cohort, endpoint and model. The Foundation would supply a validated common substrate.
Original FASTQ, BAM, reports, spreadsheets, DICOM and raw instrument files with checksums, source metadata and access controls.
Canonical alignments, calls, expression, isoforms, proteins, cells and imaging features generated with versioned pipelines.
Cross-patient and within-patient tables, common vocabularies, treatment timelines, outcome labels, QC and full provenance.
The high-value work begins before an operation: coordinate surgery, pathology, the biobank and assay laboratories so each lesion remains distinct and pre-analytic details are recorded at the source.
Connect preoperative imaging, operative location, specimen label, pathology block and molecular files through one durable lesion ID.
Use comparable DNA, short-read RNA and pathology assays at every major disease state so serial change can be measured directly.
Where tissue quantity and quality allow, add WGS, long-read RNA, single-cell and spatial, protein, immunopeptidome and functional models.
Capture exact prior dose, last exposure, drug holds, surgery timing and concurrent therapies, not only "pretreated" or "post-treated."
Diagnosis, tumor content, necrosis, IHC, whole-slide image and future spatial assays.
WGS/WES, short-read RNA, long-read RNA, proteome and phosphoproteome.
Single-cell or multiome analysis, organoids and functional drug testing where feasible.
Germline reference, sample identity, HLA and immune analysis and normal comparator.
Pre- and postoperative ctDNA or other blood biomarkers linked to imaging and treatment.
Collection time, cold ischemia, preservation, aliquot, storage and transfer conditions.
An authorized researcher should be able to open a notebook and analyze across patients, timepoints, modalities and treatments without first reconstructing what each file means.
Clean once. Align once. Validate once.
Analyze repeatedly, by patients, researchers, drug developers and future care teams.
Reference architecture informed by the NCI Genomic Data Commons data model, data processing and quality framework; GA4GH Phenopackets, Variation Representation Specification and Data Use Ontology; HL7 FHIR Genomics Reporting; OHDSI OMOP; CDISC SDTM; Common Workflow Language; and NCI Biospecimen Best Practices. These are precedents, not partners. This design is not a final implementation specification and will be co-designed with pathology, bioinformatics, privacy, regulatory, clinical and patient experts.