Four ways evidence arrives
Acquisition modes, not vendors
The market research sorts every source into one of four modes. The mode decides how much provenance you get for free and how much rights work you owe.
1. Official open APIs and bulk data
Best for government, filings, research, and some registries. Usually strong provenance, but schemas and status semantics still require interpretation.
2. Licensed commercial APIs or feeds
Best for normalized private-company, funding, people, traffic, and news data. Product embedding and derived-data rights are as important as endpoint coverage.
3. Public feeds and structured pages
RSS/Atom, sitemaps, public job boards, changelogs, press rooms, conference sites, and open repositories can be monitored with clear source links.
4. Permission-aware web collection
Some valuable evidence has no API. Collection should respect robots directives, site terms, rate limits, authentication boundaries, copyright, privacy, and source-specific retention rules. "Publicly visible" does not automatically mean "licensed for product ingestion and redistribution."
The product should not depend on scraping authenticated social feeds, evading technical restrictions, or republishing full copyrighted articles. Where rights are limited, storing a title, URL, timestamp, small compliant excerpt, extracted observation, and provenance may be more appropriate than storing or displaying the full source.
Company monitoring sources
What changed at this company: where the evidence lives
The thirteen source families from the company news research. Each row pairs the best use with the limitation that should sit next to it in the product.
| Source family | Examples | Best use | Important limitation |
|---|---|---|---|
| Open feeds and site metadata | RSS, Atom, WebSub, sitemaps, JSON-LD | Company blogs, newsrooms, changelogs, podcasts | Coverage depends on publisher quality and feed completeness |
| Direct site monitoring | Polite crawl, selected-page diff, browser rendering | Pricing, team, partners, legal pages, sites without feeds | Must respect robots/terms, suppress noise, and avoid authenticated/private pages |
| Broad web/news discovery | Google Alerts input, NewsAPI, Event Registry, GDELT, licensed media-monitoring vendors | Mentions, earned coverage, syndication, multilingual discovery | Indexes differ; full-text and redistribution rights vary |
| Company/funding data | Crunchbase, PitchBook, Tracxn, Dealroom and similar licensed sources | Identity enrichment, rounds, people, acquisitions | Commercial licensing; data may lag or conflict with primary sources |
| Social APIs | X, Bluesky/AT Protocol, Mastodon, YouTube | Official posts, mentions, videos, founder activity | Volatile terms/costs; completeness and historical access vary |
| Approved Community Management integrations, user-supplied links/alerts, licensed partners | Company-page activity and professional announcements | Vetted access; generic member-post monitoring is not openly available | |
| Podcasts and newsletters | Podcast Index, RSS, publisher feeds, user-forwarded email | Founder interviews, essays, niche coverage | Transcripts may require generation or licensing; email consent matters |
| Corporate registries | SEC EDGAR, Companies House, state/national registries | Filings, officers, legal status, offerings | Jurisdiction fragmentation and entity-ID matching |
| Regulatory/open government | openFDA, clinical trials, patents/trademarks, USAspending, procurement portals | Approvals, recalls, patents, grants, contracts | Domain expertise required; absence is not evidence |
| Hiring systems | Greenhouse, Lever, other ATS endpoints, JobPosting markup | Role/function/geography changes | Openings are intent, not hires; postings can be duplicated or stale |
| Product/developer ecosystems | GitHub, package registries, Product Hunt, status pages, marketplaces | Releases, integrations, incidents, technical momentum | Public activity may not represent the core product or revenue |
| App stores and reviews | Founder-connected App Store/Play accounts, licensed providers, public listings | Releases, ratings, customer themes | Official review APIs often cover only the authenticated developer's apps |
| Web archives | Common Crawl, Internet Archive | Historical baselines, deleted pages, name/domain history | Incomplete and delayed; preservation and reuse constraints apply |
Feeds should be preferred when available. Where no feed exists, a watched page with a readable before/after diff is the fallback, with navigation, cookie banners, timestamps, and tracking parameters suppressed so "change" means a change a human would care about. How Company Pulse uses these layers →
Market intelligence sources
What is changing in its space: where the evidence lives
The thirteen source families from the market research. The right mix depends on target industries, rights, economics, geography, and whether PostMoney may show the underlying material to customers.
| Source family | Example sources | Interesting evidence | Important limitation |
|---|---|---|---|
| Private-market data | Crunchbase, Dealroom, PitchBook, CB Insights, Tracxn, Harmonic | Companies, people, rounds, investors, valuations, M&A, taxonomies, headcount | Expensive; conflicting coverage; redistribution and derived-data rights matter |
| Public-company and corporate filings | SEC EDGAR, Companies House, state registries, competition authorities | Filings, subsidiaries, financial facts, acquisitions, legal entity formation | Private-company coverage is limited; names and entities require careful matching |
| News and web discovery | GDELT, licensed news APIs, search providers, trade publications, RSS, newsletters | Events, narrative, launches, partnerships, category coverage | Copyright, archive depth, syndication, paywalls, and false matches |
| Market statistics | Census, BLS, World Bank, OECD, Eurostat, trade associations | Establishment counts, employment, prices, demographics, imports/exports | Often lagged or too broad; classification systems rarely match a startup's market |
| Digital and app intelligence | Similarweb, Semrush, Sensor Tower, data.ai, app stores | Estimated traffic, keywords, channels, rankings, downloads, engagement, reviews | Modeled estimates; weak coverage at low scale; own-app APIs differ from competitor access |
| Commerce and advertising | Retailers, marketplaces, price trackers, Meta/Google ad libraries, affiliate data | Assortment, price, rank, reviews, promotions, creative, stockouts | Site terms and page structures change; rank is not sales; geographic blind spots |
| Hiring and people | Public ATS boards, company career sites, licensed people datasets | Open roles, functions, seniority, location, executive and talent movement | Open roles are intentions, not hires; boards can be stale or incomplete |
| Technology and open source | GitHub, package registries, documentation, changelogs, BuiltWith/Wappalyzer | Releases, contributors, dependencies, integrations, technology adoption | Identity matching is hard; activity does not prove revenue or product quality |
| Research and IP | OpenAlex, Crossref, Semantic Scholar, USPTO, EPO, Lens | Papers, institutions, authors, funders, patents, citations, topic momentum | Publication and patent incentives create noise; long lags; legal interpretation may be needed |
| Healthcare and life sciences | ClinicalTrials.gov, openFDA, CMS, PubMed, WHO registries | Trials, approvals, clearances, recalls, adverse events, reimbursement, research | Domain-specific status and causality must be interpreted carefully |
| Policy and procurement | Regulations.gov, Federal Register, SAM.gov, USAspending, grants databases, state portals | Rules, comments, enforcement, solicitations, awards, grants | Jurisdiction and lifecycle state are essential; documents can be difficult to normalize |
| Customer voice | App stores, review sites, forums, public communities, issue trackers | Pain points, switching, sentiment, feature requests, product failures | Sampling bias, manipulation, identity uncertainty, and collection restrictions |
| User-provided research | Uploaded analyst reports, interview notes, expert calls, memos, spreadsheets | Proprietary judgment, private channel checks, custom market models | Access controls, licensing, confidentiality, and source provenance |
No single provider should be treated as complete. Funding dates and amounts often disagree, undisclosed rounds are common, and announcements lag legal closings. The product could show conflicting source claims and let the investor choose a canonical observation. How the Market Dossier uses these lenses →
Open APIs the research names
The endpoints the two documents actually cite
Grouped by what they answer. Each is named in the research with a link to its documentation. Where a doc says access would need to be checked, that hedge is kept.
Filings and registries
SEC EDGAR: unauthenticated JSON for filer submissions and XBRL company facts, updated throughout the day. Companies House: most filed UK company information through a REST API. PatentsView: patent and disambiguated inventor/assignee endpoints. USPTO Open Data Portal: search across patent and trademark bulk-data products. USAspending: recipient-level federal contracts and grants.
Regulation and health
Regulations.gov: federal documents, dockets, and comments. openFDA: public datasets across drugs, devices, foods, recalls, adverse events, clearances, and approvals, with the warning that the data is not validated for clinical decisions. Its device endpoints cover 510(k)s, PMAs, recalls, adverse events, registrations, and device identifiers. ClinicalTrials.gov: a modernized data API for study records.
Research and code
OpenAlex: a connected research graph of works, authors, institutions, sources, topics, and funders. GitHub Events and Releases APIs: public organization and repository activity, with GitHub's own warning that the events feed has latency and a limited recent window. Product Hunt: a GraphQL API. Atlassian Statuspage: structured incident and component resources.
Hiring
Greenhouse Job Board API: published jobs, departments, offices, descriptions, and update timestamps through unauthenticated GET endpoints. Lever Postings API and hosted job sites. Other ATS providers, company sitemaps, and structured JobPosting markup fill gaps. Public ATS endpoints are often more dependable than scraping rendered careers pages.
Feeds and archives
Atom (RFC 4287) for syndication and WebSub for pushed feed updates over webhooks. Sitemaps for URLs and modification hints. Article structured data (NewsArticle, BlogPosting) for headline, author, and dates. Common Crawl and the Wayback CDX interface as backstops for baselines and vanished evidence, not guaranteed-complete real-time sources.
Social and media
X search and filtered stream under pay-per-use access. Bluesky: an unauthenticated author-feed endpoint. Mastodon public timelines, with completeness shaped by server configuration and federation. YouTube Data API. Podcast Index. Google Alerts as a user-forwarded or RSS input, not a general official search API. NewsAPI (body content may be truncated), Event Registry, and GDELT data plus its DOC API for article search and coverage volume.
Demand and statistics
Similarweb API: estimated web traffic, engagement, channels, apps, keywords, and competitive data. Estimates, not audited metrics, and small sites may be especially noisy. Google Trends API alpha: consistently scaled search interest with roughly five years of history; access and product status would need to be checked before relying on it. SAM.gov opportunities: solicitations and award fields including amount, awardee, agency, and NAICS code. Census developer APIs, BLS, and World Bank indicators as auditable model inputs, authoritative but usually too coarse to be a TAM answer on their own.
Private-market data
Crunchbase: private-company firmographics and round-by-round funding; its full API is a licensed product. Dealroom: companies, people, rounds, valuations, talent, news, and market maps through commercial API and feed products. PitchBook, CB Insights, Tracxn, Harmonic, and regional datasets are other candidates whose coverage, licensing, embedding rights, and economics would need direct evaluation.
Commercial media-monitoring and company-data vendors could add licensed full text, historical depth, broadcast transcripts, or stronger entity resolution. Across every family, API prices, quotas, policies, and permitted storage change frequently. Why provenance fields matter more than any one endpoint →
What no source can tell you
Each family is one lens
- Absence is not evidence. Coverage is jurisdiction- and sector-specific.
- An opening is intent, not a completed hire. Removal can mean filled, canceled, or moved.
- Rank is not sales.
- Repository stars do not equal commercial adoption.
- Patents do not equal product quality. Publication volume does not equal scientific validity. Trial status does not equal regulatory approval.
- An adverse-event report is not proof that a product caused an event. A recall record may be updated.
- A proposed rule is not final law.
- One press release republished 80 times is one signal, not 80 independent signals.
- A legal or regulatory match should never be summarized as guilt, failure, or clinical meaning.
Constrained by design
Two sources the product should never assume
The research singles out two places where a universal public API does not exist, and says what the realistic inputs are instead.
Treated as a constrained source, not something PostMoney can generically scrape forever. Most LinkedIn permissions require explicit approval; Community Management is a vetted program with development and standard tiers, and member-post read access is closed.
Realistic inputs: user-supplied LinkedIn links, approved page integrations, alert emails, or licensed data, with coverage gaps clearly labeled.
App reviews
Apple's App Store Connect customer-review API and Google Play's reviews endpoint are designed for the developer's own apps and require appropriate authorization. Review-site terms may restrict automated collection or redistribution.
Rich review monitoring for arbitrary portfolio companies may therefore require a founder-connected account, a licensed provider, or carefully permitted public-page collection, not an assumption that a universal public API exists.