Clarida Foundation
The data foundation

Collect once. Clean once. Analyze many times.

Every contribution should enter the Foundation's data as a complete, longitudinal, analysis-ready patient record, not as another folder of disconnected files.

This page is a design. It describes how the Foundation intends to collect, clean, govern and share the data it will hold. None of it is running today, and the legal statements in it are drafts for counsel and IRB review.

  • All modalities connected
  • Every timepoint preserved
  • Source data never overwritten
  • Harmonized at contribution
Arrives as
FASTQ / BAM / VCFDNA, RNA, long-read
Reports / spreadsheetsLab calls and interpretation
Clinical / imagingTreatment, toxicity, outcome
Consent + use rightsFHIR Consent · DUO terms · tier · version
Harmonized
at
contribution
Leaves as
Longitudinal patient graphEvery sample and exposure linked
Analysis-ready matricesConsistent coordinates and vocabularies
Queryable, with the consent attachedAsk the cohort question first
No file hunting. No manual sample matching. No coordinate liftover. No bespoke ETL.
~250files in one illustrative longitudinal patient catalog
~559 GBraw and lab-processed molecular data
5source folders with different lab conventions
5+serial tissue timepoints across treatment states
7+molecular and functional modalities represented or planned
Private patient identifiers and exact dates are omitted from this public page.
01 / Complete patient record
What data is critical

The biological data only become useful when they remain attached to the patient journey.

The unit of analysis is not a sequence file. It is a patient, a disease state, a lesion, a specimen, a treatment exposure, and an outcome, linked over time.

Clinical timeline + treatment exposure
Imaging, pathology + lesion response
ctDNA, labs, symptoms + outcomes
DNA: WGS, WES, targeted panels
The patient one continuous record
RNA: short-read + long-read
Proteome + phosphoproteome
Single-cell + spatial biology
Immunopeptidome + immune state

Clinical context

Diagnosis, disease burden, performance, comorbidities, every treatment, dose, hold, surgery, radiation, response and toxicity.

Biospecimen context

Lesion identity, anatomic site, collection time, preservation, tumor content, necrosis, cold ischemia and pathology review.

Raw molecular data

Instrument-level files whenever available, retained immutably so future methods can reprocess the original evidence.

Lab-derived results

Original reports, spreadsheets, calls and interpretations are preserved beside the harmonized analysis, not substituted for it.

Longitudinal outcomes

Early molecular response, imaging, symptoms, quality of life, progression, resistance and subsequent interventions.

Complete provenance

Who generated each file, with which instrument, reference, software, parameters, annotation set and quality status.

02 / Complementary sequencing
Different tests answer different questions

WGS, WES, RNA, long-read and single-cell data are not interchangeable.

The record should preserve the strengths and limitations of each assay, then connect them into a coherent view rather than flattening them into one vendor report.

DNA / WGS

Whole genome

Broadest view of inherited and tumor-acquired change.

  • SNVs and indels
  • Copy number and structural variation
  • Non-coding and breakpoint context
  • Best paired with matched normal
DNA / WES

Whole exome

Efficient depth across protein-coding genes and a strong cross-validation layer.

  • Actionable coding variants
  • Tumor-normal comparison
  • Serial clonal change
  • Capture-kit effects retained
DNA / Targeted

Clinical panels

Useful clinical snapshots with fast interpretation, but limited genomic breadth.

  • Hotspots and selected genes
  • Clinical lab assertions
  • Panel-covered negatives only
  • Never treated as equivalent to WGS
RNA / Short-read

Transcriptome

Shows which genes and pathways are active at a particular disease state.

  • Expression counts and TPM
  • Fusion support
  • Immune and pathway signals
  • Allele-specific expression
RNA / Long-read

Full-length isoforms

Resolves transcript structure that short reads can fragment or misassign.

  • Isoforms and splice junctions
  • Fusion transcript architecture
  • Novel transcript discovery
  • Cross-check against short-read RNA
Cells / Spatial

Heterogeneity

Separates tumor, immune and stromal states that bulk measurements average together.

  • Cell populations and states
  • Rare resistant subclones
  • Cell-cell interactions
  • Spatial organization when available

One timepoint is a snapshot. Repeated sequencing turns the disease into a movie.

03 / Longitudinal by design
Every specimen is anchored to treatment

Track what changed, when it changed, and what pressure selected it.

Serial tissue and blood allow within-patient comparisons across diagnosis, treatment, recurrence and metastasis. Multiple lesions collected at the same procedure reveal spatial heterogeneity as well as temporal evolution. The timepoints below are from the illustrative catalog.

  1. 01
    T0

    Pre-treatment biopsy

    Baseline biology before systemic therapy.

    WESWTSpanel
  2. 02
    T1

    Resection after initial therapy

    Residual disease after first treatment pressure.

    WESWTSpathology
  3. 03
    T2

    Recurrence biopsy

    Biology after completion of standard therapy.

    WESWTSreport
  4. 04
    T3

    Resection after salvage exposure

    Fresh-frozen and FFPE material for deeper analysis.

    WGS*WESWTSlong RNAprotein
  5. A+B
    T4

    Two lung metastases

    Independent nodules with shared matched normal.

    WESWTSproteinimmune
  6. next
    T5

    The next resection

    Repeat the core assays and add high-value modalities where tissue permits.

    WGSWTSlong RNAsingle-cellfunctional
Treatment exposures, dose changes, holds, surgery and radiation are represented as timed events between every sample.

*An expected WGS component in the illustrative catalog was not present in the delivered folder. The record would mark it expected-but-missing, not as a negative result.

04 / One sample lineage
A graph, not a file cabinet

Every result must be traceable back to the exact patient, lesion, specimen and assay.

Stable identifiers and validated relationships prevent the most damaging class of errors: mispaired normals, merged lesions, ambiguous timepoints, lost aliquot history and results detached from treatment context.

Patient
Disease episode
Procedure
Lesion
Specimen
Aliquot / analyte
Assay run
Raw file
Harmonized file
Derived feature

Every node has a durable ID.

The graph keeps biological replicates, technical replicates, multiple nodules, shared normals and repeated assays distinct while making their relationships directly queryable.

A file without lineage is stored, but it is not labeled analysis-ready.
05 / Clean at contribution
The work researchers should not have to repeat

Ingest, validate, align, normalize and document the data once.

A contribution would count as complete only after it passes a data contract and produces versioned, quality-scored, analysis-ready outputs.

  1. 1

    Manifest + integrity

    Inventory every expected and received file; verify checksums, byte size, compression and readability.

  2. 2

    Identity + pairing

    Fingerprint samples, confirm patient identity, tumor-normal pairing, lesion IDs, sex and contamination.

  3. 3

    Metadata contract

    Map site, timepoint, preservation, assay, platform, panel, read group and treatment context to controlled terms.

  4. 4

    Technical QC

    Coverage, duplication, insert size, mapping, RNA quality, tumor purity, read structure and modality-specific checks.

  5. 5

    Canonical alignment

    Reprocess DNA and RNA against a versioned common reference and annotation bundle while retaining source alignments.

  6. 6

    Standard calling

    Generate SNV/indel, CNV, SV, fusion, expression, isoform and other derived outputs with reproducible workflows.

  7. 7

    Normalize + annotate

    Normalize coordinates and identifiers; add transcript, gene, protein, pathway and clinical annotations with versions.

  8. 8

    Reconcile platforms

    Measure concordance, retain disagreement, identify report-only evidence and never convert missing into negative.

  9. 9

    Privacy + use rights

    Apply de-identification, access tier, consent and data-use tags. Preserve intervals in longitudinal research views. Section 06 describes the machine-readable stewardship layer.

  10. 10

    Publish + refresh

    Release immutable raw, harmonized and analysis-ready layers through versioned snapshots.

Modality-specific harmonization
Data typeWork completed at contributionResearch-ready result
DNACommon reference; read-group validation; tumor-normal identity; SNV/indel, copy number and structural-variant processing; variant normalization.Comparable variant and copy-number tables across laboratories, panels and timepoints, plus source-native calls.
Short-read RNACommon genome and transcript annotation; gene-level counts; expression normalization; fusion analysis; quality and batch covariates.Counts, TPM, fusions, pathway-ready matrices and explicit assay/batch metadata.
Long-read RNAHiFi QC; transcript alignment; isoform collapse and classification; junction and fusion reconciliation with short reads.Full-length isoform and transcript-event tables linked to expression and genomic breakpoints.
Single-cell / spatialBarcode and UMI QC; doublet and ambient-RNA assessment; raw-count preservation; cell and region annotations; batch flags.Portable matrices with curated cell labels, embeddings, QC fields and spatial coordinates where available.
Protein / peptideStandard protein and phosphosite IDs; FDR and intensity normalization; peptide-to-protein and variant-peptide mapping; raw-spectra status.Comparable protein, phosphosite and immunopeptide features with confidence and provenance.
Clinical / imagingMedication, lab, disease, toxicity and response vocabularies; lesion IDs; DICOM de-identification; treatment intervals and outcomes.Queryable treatment-exposure, lesion-response, toxicity and outcome tables synchronized to molecular timepoints.
07 / What one real catalog teaches
Heterogeneous inputs are normal

The design must understand incomplete, overlapping and differently processed contributions.

The illustrative catalog includes raw reads, aligned files, reports, spreadsheets, multiple laboratories, FFPE and flash-frozen tissue, shared normals, tumor-only data, report-only deliveries and expected files that have not yet arrived.

Illustrative de-identified source catalog

What comes in, and what would come out
Multi-timepoint WES + WTSFFPE · multiple deliveries · raw + processed + reports
Five serial analyses, but one timepoint is report-only and another delivery contains FASTQ without BAM, spreadsheets or checksums.
Reconcile orders and specimens; canonical WES/WTS reprocessing where raw data exist; source-report extraction; explicit raw-data availability and QC fields.
Output: serial DNA/RNA matrices
Two independent lung nodulesFlash-frozen · WES + WTS · one shared normal
Separate tumor folders use a single germline reference. The relationship is valid but easy to misread without explicit graph metadata.
Confirm fingerprint and normal identity; assign durable lesion A/B IDs; process both tumors against the same normal; enable within-lung heterogeneity analysis.
Output: lesion-resolved comparison
PacBio long-read RNAHiFi BAM + Iso-Seq outputs
The folder name references tumor/PBMC WGS and long-read RNA, but only the long-read delivery is present.
Inventory expected versus received modalities; publish the isoform layer; keep tumor and blood WGS as missing/awaited, not absent biology.
Flag: expected WGS missing
Tumor-only DNA + RNAClinical profiling · no matched normal
Paired DNA and RNA reads are available, but inherited variation cannot be filtered with a patient-matched normal from that assay.
Preserve clinical calls; run tumor-only processing with explicit uncertainty; cross-check against the patient's normal from other sources where consent and methods allow.
Flag: tumor-only interpretation
Protein-level reportsPhosphoproteome PDF + abundance spreadsheet
Processed results exist without the same raw-data depth as the sequencing folders.
Map protein and phosphosite identifiers; retain original z-score reference and report provenance; label results as report-derived if raw spectra are unavailable.
Flag: report-derived evidence
Planned / pending assaysImmunopeptidomics · single-cell · organoids · phosphoproteomics
Some tests are planned, pending, performed elsewhere or not yet delivered.
Use status values (planned, collected, processing, delivered, failed QC, unavailable) so absence of data is never mistaken for a negative assay result.
Output: complete audit trail
Missing is not negative.
Discordance is measured, not erased.
Raw and source-native results remain available.
08 / Research-ready by default
The researcher experience

Analysis should begin with a cohort question, not months of reconstruction.

"No prep work" should mean no foundational wrangling. Researchers still choose the scientific question, cohort, endpoint and model. The Foundation would supply a validated common substrate.

01

Immutable source layer

Original FASTQ, BAM, reports, spreadsheets, DICOM and raw instrument files with checksums, source metadata and access controls.

02

Harmonized molecular layer

Canonical alignments, calls, expression, isoforms, proteins, cells and imaging features generated with versioned pipelines.

03

Analysis-ready patient layer

Cross-patient and within-patient tables, common vocabularies, treatment timelines, outcome labels, QC and full provenance.

Example cohort query
Find metastatic sarcoma patients with a recurrent fusion event; compare DNA support, short- and long-read RNA, pathway activity and protein signaling before and after pathway-directed treatment; relate the changes to ctDNA, lesion response and toxicity.
same referencesame IDssame time axisquality filteredprovenance attached
  • patient_timeline.parquet
  • specimens.parquet
  • variants_long.parquet
  • copy_number.parquet
  • structural_variants.parquet
  • expression.parquet
  • fusions_isoforms.parquet
  • single_cell.h5ad
  • proteomics.parquet
  • treatment_exposure.parquet
  • lesion_response.parquet
  • provenance.json
Researchers should never have to spend their first six months reassembling the patient.
09 / Use new tissue to set the standard
When new tissue is collected

Capture each nodule once. Preserve enough material to answer future questions.

The high-value work begins before an operation: coordinate surgery, pathology, the biobank and assay laboratories so each lesion remains distinct and pre-analytic details are recorded at the source.

Map every lesion

Connect preoperative imaging, operative location, specimen label, pathology block and molecular files through one durable lesion ID.

Repeat a core panel

Use comparable DNA, short-read RNA and pathology assays at every major disease state so serial change can be measured directly.

Add deeper modalities

Where tissue quantity and quality allow, add WGS, long-read RNA, single-cell and spatial, protein, immunopeptidome and functional models.

Record the treatment state

Capture exact prior dose, last exposure, drug holds, surgery timing and concurrent therapies, not only "pretreated" or "post-treated."

Each nodule kept separate
FFPE + digital pathology

Diagnosis, tumor content, necrosis, IHC, whole-slide image and future spatial assays.

Flash-frozen tissue

WGS/WES, short-read RNA, long-read RNA, proteome and phosphoproteome.

Viable cryopreserved cells

Single-cell or multiome analysis, organoids and functional drug testing where feasible.

Matched blood / PBMC

Germline reference, sample identity, HLA and immune analysis and normal comparator.

Serial plasma

Pre- and postoperative ctDNA or other blood biomarkers linked to imaging and treatment.

Pre-analytic metadata

Collection time, cold ischemia, preservation, aliquot, storage and transfer conditions.

Clinical diagnosis and required pathology always take priority. Research allocation must follow a pre-approved collection plan.
The data-readiness commitment

The research begins where data wrangling ends.

An authorized researcher should be able to open a notebook and analyze across patients, timepoints, modalities and treatments without first reconstructing what each file means.

  1. 01Every source file is immutable and checksum-verified.
  2. 02Every specimen and derived result has traceable lineage.
  3. 03Every harmonized output is versioned and reproducible.
  4. 04Every missing or uncertain element is explicit.
  5. 05Every new timepoint refreshes the longitudinal record.
  6. 06Every record carries the patient's consent, machine-readable and enforced at query time.

Clean once. Align once. Validate once.
Analyze repeatedly, by patients, researchers, drug developers and future care teams.

  • NCI GDC-inspired graph + harmonizationCase, biospecimen, assay and file lineage; validation; common processing.
  • GA4GHVRS, Phenopackets, data-use terms and portable, federated workflow interfaces.
  • HL7 FHIR + OMOPComputable genomic reporting and standardized longitudinal clinical data.
  • CDISC + DICOMRegulatory trial datasets and interoperable imaging with structured lesion tracking.
  • CWL / containersPortable, reproducible bioinformatics workflows with explicit inputs and outputs.
  • BioCompute provenanceHuman- and machine-readable workflow versions, parameters and execution history.
  • Controlled vocabulariesGene, variant, disease, drug, lab, toxicity, anatomy and outcome concepts mapped once.
  • Governed accessPublic aggregate knowledge; controlled individual-level data; secure computation when needed.
  • DUO + FHIR ConsentMachine-readable permitted-use terms attached to a versioned patient consent resource.
  • GA4GH PassportsVerified researcher identity and access claims used by the policy check.
  • Single IRB relianceA master addendum and reliance model for cooperative research where applicable.
  • Common Rule + HIPAABroad consent, authorization, access, de-identification and withdrawal rules.
  • GDPR + transfer controlsResearch basis, safeguards, data-subject rights and an approved international transfer mechanism.

Reference architecture informed by the NCI Genomic Data Commons data model, data processing and quality framework; GA4GH Phenopackets, Variation Representation Specification and Data Use Ontology; HL7 FHIR Genomics Reporting; OHDSI OMOP; CDISC SDTM; Common Workflow Language; and NCI Biospecimen Best Practices. These are precedents, not partners. This design is not a final implementation specification and will be co-designed with pathology, bioinformatics, privacy, regulatory, clinical and patient experts.