M
Made Scientific
Commercial AI-OS · built with BioCreative

Data & Systems Inputs Map

Every database, external API, CRM, and provider feeding Made's commercial engine — 8 live Supabase databases and 26 input channels, grouped into 5 swimlanes. These are parallel, independent sources, not a sequential pipeline. Several feed directly into the stages on the companion Account & Contact Lifecycle page. Click any source to jump to its reference card.

Live-verified 2026-07-28 — every row count, cron schedule and Edge Function checked against the running databases and the bc-made VPS, not against the written docs
Companion page → Made — Account & Contact Lifecycle · how a row gets classified, enriched, identity-resolved and scored
Core store — the engine lives here Live & syncing on a schedule Live but manual / stale mirror Gap — built but not flowing Designed or retired Free — public API or internal SQL Paid — licensed data, LLM, or metered API
The estate

Eight live databases, one anchor

Made's data is not one database — it is a federation. Exactly one of them, Transfer / MAID, is where the commercial engine actually runs; everything else either feeds it, mirrors it, or serves a separate function entirely. Knowing which is which is the difference between trusting a number and double-counting it.
Swimlane 1 · Upstream

BioCreative Hub → MAID

The classification upstream, plus eleven scheduled life-science mirrors that keep Made's copy of the market current without ever re-hitting a public API.
Swimlane 2 · CRM & business systems

The second database — Reporting, and what flows both ways

This is the layer that makes the engine a real commercial system rather than a marketing database: Salesforce and HubSpot are the systems of record Made's team actually works in. Most of it is read-only into the engine — but Monday is genuinely bi-directional.
Swimlane 3 · Enrichment & research providers

The paid engines — and who pays

Five composable providers. No single vendor does the whole job, so they are orchestrated: Clay answers who works there and what the company is; FullEnrich and Brave answer what is their domain, email, and LinkedIn; Firecrawl answers what does their own website say. All Clay and FullEnrich spend sits on Made's own accounts.
Swimlane 4 · Life-science & market intelligence

The domain corpus that becomes a signal

All of these are free public APIs except GlobalData, which is licensed. They matter here for one reason: they are the raw material of the dynamic momentum score — but only for accounts they can be linked to.

Exhibit — how a company's modality is actually decided

Made's most important scientific field is not set by one classifier. It is resolved by a 17-source fallback cascade, tried in rough order of trust, each stamping its own name into modality_classification_source. Live counts, 2026-07-28. Read this as the answer to “where did that modality come from?” — and note what is absent.

1,907
hub_ai_modality — BioCreative Hub's AI classification. The dominant source by a wide margin.
681
hub_sync — inherited wholesale from the Hub record.
285
clinical_trials_molecule_type — read off the molecule type of the company's registered trials.
163
ai_finder — the discovery agent's own classification at intake.
133
ct_studies_text — inferred from free-text ClinicalTrials.gov study descriptions.
83
description_text — inferred from the company's own marketing copy.
66
gd_modality_source — licensed GlobalData drug-pipeline modality, rolled up by update_company_modality_from_drugs().
96
ten weaker sources combined drug_type 26 · ctgov_api 21 · drug_type_vaccine 17 · drug_type_small_molecule 16 · description_keyword 12 · description_text_broad 9 · drug_type_gene_therapy 7 · hub_ai_subcategory 4 · drug_type_protein 1.
1,164
No source at all — and no modality — a quarter of the estate never reached any rung of the ladder.
287
allo_auto resolved — 6.2% — the single most commercially important CGT distinction, and the cascade almost never produces it.
117
manufacturing_modality — 2.5% — what the account actually needs manufactured.

What is absent from all 17 rungs: Salesforce. sfdc_accounts.account_modalities carries rep-curated subtypes (CAR-T, TCR-T, MSC, iPSC, CAR-NK, HSC, Treg, TIL) for 781 accounts and appears nowhere in this cascade. Those labels strongly imply allo/auto — the very field sitting at 6.2%. With 1,268 SFDC accounts already MADE-ID-linked, this is a join, not a research project.

Also worth knowing: nine different modality columns are populated on companies with no documented canonical (modality_category 3,043 · therapy_category 2,561 · cell_modality_primary 2,299 · ai_modality 2,241 · technology_platform 2,224 · primary_indication 1,950 · platform_type 529 · modality_ai_classification 163 · modality_type 120), while three the documentation lists as filled hold zero rows (pipeline_category, primary_therapeutic_area, therapy_types). Different app surfaces reading different columns is why modality counts disagree between screens.

Swimlane 5 · Channel feedback

What the outbound channels write back

Included because it closes the loop: a reply or a LinkedIn accept is itself an input. It is what flips an account's engagement flags, and therefore what the relationship ladder reads.
The one thing to understand about this map

Collection is solved. Linkage is the constraint.

Nearly every channel on this page is live and current — 38 scheduled jobs on Transfer alone, eleven life-science mirrors, both CRMs syncing daily, news arriving every four hours. The engine is not short of data.

What it is short of is attachment. A funding announcement only moves an account's score if it is linked to that account; a HubSpot contact only enriches a company if it resolves to that company. That is why company_match_queue at 9,044 rows and made_id_contact_assign_queue at 40,531 rows are the two most consequential numbers in the whole estate — and why draining them, not adding another source, is the highest-leverage work available.

Full reference

What each source captures, where it lands, and what it costs

Every card names the real Edge Function, sync function, cron entry, and landing table, with live row counts pulled 2026-07-28. Where a figure contradicts wiki/reference/data-lineage.md — which was last accurate in March 2026 — the card says so.

The database estate
Swimlane 1 — BioCreative Hub → MAID
Swimlane 2 — CRM & business systems
Swimlane 3 — Enrichment & research providers
Swimlane 4 — Life-science & market intelligence
Swimlane 5 — Channel feedback