Data Methodology

How we collect, normalize, and version peptide therapeutic data.

Peptide Pipeline is built on a normalized, API-agnostic content repository. The data source can change without rewriting page components, and every material assertion is tied to a controlling source.

Primary-source stack

  • ClinicalTrials.gov study records and API downloads
  • PubMed / NCBI E-utilities
  • Drugs@FDA and openFDA endpoints
  • FDA approval letters, labels, and advisory records
  • EMA medicine and document data
  • SEC EDGAR APIs and filings
  • Company investor-relations filings and press releases
  • Conference abstracts and presentations
  • Patent-office records

Normalization

We normalize stages, indications, targets, modalities, and sponsor names into controlled vocabularies. Aliases are preserved and linked to canonical records. Assets are not merged solely on target, sponsor, similar code, or claimed sequence.

Versioning

Stage assertions, catalysts, and deal records are versioned. Each carries a valid-from date, reviewer, controlling source, and confidence level. Superseded assertions remain in change logs.