Data Methodology

How we collect, normalize, and version peptide therapeutic data.

Primary sources

  • ClinicalTrials.gov study records and API downloads
  • PubMed and NCBI E-utilities for peer-reviewed literature
  • Drugs@FDA data files and openFDA endpoints
  • FDA approval letters, labels, and advisory records
  • EMA medicine and document data
  • SEC EDGAR APIs and company investor-relations filings
  • Conference organizers’ official abstracts and presentations
  • Patent-office records such as USPTO, EPO, and WIPO

Normalization

We normalize development stages, indications, targets, modalities, and company names into a controlled vocabulary. Aliases are preserved and linked to canonical records. We do not merge assets solely on target, sponsor, similar code, or claimed sequence.

Versioning

Stage assertions, catalysts, and deal records are versioned. Each assertion carries a valid-from date, reviewer, controlling source, and confidence level. Superseded assertions remain visible in timelines and change logs.

Review cadence

  • High-interest assets: at least every 90 days
  • All active clinical assets: at least every 180 days
  • Catalyst records: monthly or when source guidance changes