Methodology & data sources

wafergraph is independent, with no paid placements. Financials (revenue, R&D, net income, growth) come from SEC EDGAR XBRL filings; company facts (founding year, employees, HQ, website) from Wikidata; legal entity and HQ data from GLEIF; company descriptions from Wikipedia; and market-share figures are widely-reported consensus estimates shown with their year, basis, and source, grounded against each company's real current revenue. Supply-chain edges — the key_supplier and key_customer relationships behind every chokepoint and dependency figure — are curated, not scraped, and they come in two kinds. Some are researched company by company from public disclosure (company filings, product documentation, and public supplier or customer announcements) and re-verified before entry. Others are INFERRED FROM SEGMENT MEMBERSHIP: one customer list applied to every company in a segment, on the reasoning that a specialty-materials supplier really does sell to the major fabs. That reasoning holds in aggregate and it is not individual research, so here is the split, recomputed from the dataset on every build. An edge set counts as segment-inferred when its exact membership is shared verbatim by 5 or more companies. Of the 462 companies carrying a customer list, 70 lists are unique to one company (individually researched), 49 are shared by two to four companies, and 280 — about 61 percent — are segment-inferred, drawn from just 16 distinct templates; the largest single template covers 78 companies all carrying the identical list GlobalFoundries, Intel, Micron Technology, Samsung Electronics, SK hynix, SMIC, TSMC. A further 63 of those customer lists are a full segment template plus a handful of individually-added names (about 14 percent) — counted with the segment-inferred bucket above, not the researched one, because the template does most of the work. On the supplier side, 97 of 324 are unique, 62 shared by two to four, 3 template-plus, and 162 segment-inferred. What that means if you are using the shared-upstream analysis: when two vendors surface as sharing an upstream supplier, check which kind of edge produced it. Two individually researched edges are a finding. If either side is segment-inferred, it is a well-founded hypothesis about how that segment generally sells — worth confirming with the vendor, not worth acting on unconfirmed. A segment template also cannot know what a specific company actually makes, which is how a supplier of EUV-only equipment ended up listed against fabs that run no EUV; those are corrected as they are found. Per-company provenance labels for every edge set are published at https://wafergraph.com/data/edge_provenance.json, and the rule that produces them is a script in the repository rather than a judgement call. Coverage is deliberately partial and every in-degree count is a LOWER BOUND, not an exhaustive census. Three consequences worth stating plainly. First, a company's chokepoint count can rise as edge coverage improves rather than because the industry changed (193 verified edges were added on 2026-07-19, which moved TSMC's in-degree from 83 to 98). Second, because the tracked graph has a densely connected core, transitive "downstream reach" saturates — most major hubs reach a similar number of companies, so reach is a poor discriminator between them and direct in-degree is the more meaningful figure. Third, and most important if you are using the chokepoint ranking to make a decision: in-degree currently reflects how thoroughly a company's dependents have been DOCUMENTED here, not only how critical it is. As of this build, the median in-degree of companies wafergraph classifies as a market monopoly is 3, and 59 of 112 monopoly- or leader-classified companies have three or fewer recorded dependents — several have none, despite being depended on across the industry in reality — while TSMC sits at 98. A heavily documented hub therefore outranks a structurally more critical but less documented supplier. Read the ranking as "most-documented dependencies", treat a low count on a monopoly-classified company as a coverage gap rather than evidence of low criticality, and weigh the market-position field alongside it. Closing that gap is the active priority for the dataset. Fourth, on citations: an automated pass over 151 annual filings (of 163 tracked companies that file with the SEC) resolved 9 supplier relationships to a specific quotable disclosure, about 0.6 percent of the 1470 recorded dependencies. Only those 9 carry a machine-checkable citation attached to the record; the rest rest on the curation described above — individual research for some, segment inference for the majority — and most of the chain's critical suppliers do not file with the SEC at all. Fifth, 71 of those supply edges carry a verbatim primary-source citation curated by hand — the exact sentence, the filing or document it came from, and a link to open it — visible next to a company's suppliers or customers on any company page. Segment counts: the taxonomy has 12 top-level segments; market-structure concentration (CR1/CR3/CR4, HHI) is computed for the 17 sub-segments that have named market-share data, which is why both numbers appear on the site. Data last refreshed 2026-09-05.