Public Reference

Industry Primers

Bottom-up NAICS industry primers written for both public-market and private investors. Leaf industries are researched from the ground up; every group, subsector, and sector above them reads as a contrast across the industries beneath it.

2122 industries · 24 sectors · NAICS 2022

Researched with AI assistance from official U.S. statistics and independent sources, with citations on every page. Figures are not individually verified against pinned evidence — primers marked Evidence-verified are. Industry research, not investment advice. Methodology.

SubsectorNAICS 513Information

Publishing Industries (U.S.) — NAICS 513

A Histometrics rollup primer for public-market and private investors. This is a NAICS subsector (3-digit) that gathers two child industry groups. It synthesizes the two already-written child primers plus our ground-truth federal statistics for this level. Figures with citations are reported facts; statements about the future are labeled as judgments. Federal statistics are the ground truth; company and third-party figures are flagged as estimates.

NAICS = North American Industry Classification System, the U.S. federal standard for sorting businesses by activity. A 3-digit code is a "subsector"; the 4-digit codes beneath it are "industry groups," the 5-digit codes are "NAICS industries," and the 6-digit codes are "national industries."


1. Overview

NAICS 513 is where the U.S. statistical system files every business whose product is intellectual property (IP) that is published and licensed to an audience — whether that IP is words and pictures or computer code. It contains two children: 5131 (Newspaper, Periodical, Book, and Directory Publishers — "traditional" or content publishing) and 5132 (Software Publishers). The 2022 NAICS revision deliberately gathered both kinds of publishing under this one roof; before then, software publishing was numbered 511210 and content publishing sat in 5111, both inside the old 511 "Information" subsector [3][6].

The logic that unites them is a business model, not a product: in each child the publisher owns or licenses the underlying IP, bears the risk of creating it, and sells access to it — while the cost of reproducing each additional copy is near zero. A newspaper archive, a scholarly journal, a book backlist, and a piece of enterprise software are all the same kind of asset in this respect: build it once, license it many times. That shared logic is why "Publishing Industries" is a sensible statistical family.

But the two children could hardly be more different in size, growth, and investability, and that contrast is the single most important thing to understand about this level. Federal data put the whole subsector at roughly $592.3 billion of receipts in 2022 across about 30,116 firms [1] — but software (5132) is about five-sixths of that money and traditional publishing (5131) is the remaining sixth [2][3]. One child is one of the most valuable, fastest-growing, and most public-market-accessible industries in the economy; the other is a large, mature, mostly-private slice of media in structural transition. The subsector average tells you almost nothing — the only useful analysis is child-by-child, and that is the subject of this page.

Two forces run through both children and are worth holding in mind:

  • Own-the-IP, license-the-access economics. Everywhere in 513, the prize is a durable, repeating revenue stream — a subscription, a site license, a backlist, a software renewal — sold off an asset that costs almost nothing to reproduce. Scale, retention, and rights ownership decide returns in both worlds.
  • Artificial intelligence (AI) as the shared swing factor. For software, AI is both a new product to sell and a driver of usage; for traditional publishing, AI is simultaneously a new customer (licensing archives to train models) and a new competitor ("zero-click" answer engines that satisfy readers without a visit). AI now touches the economics, regulation, and risk profile of every business in the subsector [3][8].

2. What's inside — the two child industries and how they differ

NAICS 513 splits into two industry groups. They share a publishing logic but sit in different economic universes. Read this contrast before any subsector-level number.

Child (4-digit) Receipts (share of 513) Firms (share) Revenue per firm Concentration (CR4) Direction of travel Who owns them How to invest
5131 — Newspaper, Periodical, Book & Directory Publishers (content) ~$99.4B (~17%) [2] 13,326 (~44%) [2] ~$7.5M [2] CR4 14.2%; HHI 110 [2] Mature, mostly declining, internally bifurcating — commodity print erodes; a few recurring-revenue pockets (scholarly journals, book backlist, proprietary data, greeting cards) thrive Mostly private or foreign; clean listed pure-plays exist only in newspapers and scholarly-information houses; everything else is a segment inside a diversified owner Pick the segment; scarce public pure-plays, richer private opportunity set
5132 — Software Publishers ~$492.9B (~83%) [3] 16,824 (~56%) [3] ~$29.3M [3] CR4 24.3%; HHI suppressed [3] Growing, structurally attractive — recurring-subscription (SaaS) shift; AI a new product and a usage driver Unusually public-market-accessible (Microsoft, Oracle, Salesforce, Adobe, Intuit) plus a deep private-equity and venture layer Deep, liquid public menu (individual stocks + sector ETFs); also large PE/VC private routes

(CR4 = the combined revenue share of the four largest firms; CR8/CR20/CR50 extend that to the top 8, 20, and 50. HHI = Herfindahl-Hirschman Index, the sum of every firm's squared market share, which U.S. antitrust agencies treat as "unconcentrated" below 1,500 [7]. SaaS = software-as-a-service, software rented as a recurring subscription rather than bought once. PE = private equity; VC = venture capital; ETF = exchange-traded fund, a basket of stocks that trades like a single share.)

How to read the contrast. Four things stand out:

  1. The subsector is really a software industry with a traditional-publishing minority attached. Software is ~83% of receipts against traditional publishing's ~17% [2][3]. Any statement about "publishing" at the 513 level is, by weight, mostly a statement about software — so never let the content-publishing story (which is what most people picture) stand in for the number.
  2. Firm counts are far closer than revenue. Traditional publishing has ~44% of the firms but only ~17% of the money; software has ~56% of firms and ~83% of the money [2][3]. That gap is the whole story of revenue per firm: a software publisher averages ~$29 million of receipts versus a content publisher's ~$7.5 million [2][3] — and even that four-to-one gap understates it, because software's revenue concentrates in a handful of giants earning tens of billions each.
  3. Opposite directions of travel. Software demand broadly rises with automation and AI; traditional publishing is mature and, in its commodity print segments, in secular decline [3][2]. The two children are moving apart, not together — which is exactly why the blended average is misleading.
  4. Ownership is the practical dividing line. Software is one of the most investable industries in public markets; traditional publishing is one of the least, with durable value sitting mostly in private and foreign hands [3][2]. Where you can buy the industry, and how, flips completely between the two children.

3. Size (this level's rollup figures)

Our ground-truth federal file for NAICS 513 reports the 2022 Economic Census figures below [1].

Metric (NAICS 513, whole subsector) Value Source / year
Industry receipts (revenue) $592.3 billion ($592,340,719 thousand) Economic Census, 2022 [1]
Firms 30,116 Economic Census, 2022 [1]
Four-firm concentration (CR4) 20.2% Economic Census, 2022 [1]
Eight-firm concentration (CR8) 26.7% Economic Census, 2022 [1]
Twenty-firm concentration (CR20) 35.6% Economic Census, 2022 [1]
Fifty-firm concentration (CR50) 48.2% Economic Census, 2022 [1]
Herfindahl-Hirschman Index (HHI) Suppressed in federal data — not stated here Economic Census, 2022 [1]

Our ground-truth file for this level carries receipts, firm count, and concentration ratios only. The HHI is suppressed in the federal source, so no subsector HHI is stated (5131's own HHI is 110; 5132's is likewise suppressed — the two cannot be combined into a level figure) [1][2][3]. The file has no level-wide employment, payroll, or establishment line, so none is stated here as a subsector total.

How the subsector breaks down. The two children sum almost exactly to the total. Receipts: $99.4B (content) + $492.9B (software) ≈ $592.3 billion [2][3]. Firm counts sum to 30,150 against the reported 30,116 — the ~34-firm gap is the ordinary result of a few firms operating in both children and being counted once at the subsector level but in each child below. For readers who want a workforce picture, the child primers' County Business Patterns (CBP) 2023 figures add to roughly 1.26 million paid workers (≈250,000 in content publishing + ≈1.01 million in software) and on the order of $224 billion of annual payroll — but these are assembled from the children and are not part of this level's ground-truth file [5].

Why the level concentration statistic is a weak competition gauge. The subsector CR4 of 20.2% [1] sits between the two children (content's 14.2% and software's 24.3%) [2][3] — which makes sense once you see that the four largest firms of the whole subsector are almost certainly the largest software publishers, and that pooling two distinct product markets dilutes any single firm's share. A magazine does not compete with an enterprise database, and a book does not compete with a video game; when you pool them into one $592.3 billion denominator, every firm's measured share shrinks below its share of its own market. Real market power lives at the child level and below it — in the software oligopolies of specific categories, in local newspaper monopolies, in the scholarly-journal and greeting-card duopolies. Never read competition off the 3-digit number.

Undercount caveat (material, and in two opposite directions). The Economic Census counts mainly employer firms with payroll, and both children leak large amounts of real activity out of that frame — but for different reasons:

  • Traditional publishing hides a giant nonemployer tail. More than 2.6 million self-published book titles carried an ISBN (International Standard Book Number) in 2023 alone, most from sole-proprietor authors who never appear in the employer count; magazine and directory publishing hide long tails of freelancers, solo list brokers, and one-person newsletters; and the "other publishers" niche is numerically dominated by solo card artists and hobby sellers [2]. Where small or individual ownership dominates, the count runs low.
  • Software hides output that is reclassified elsewhere. A large share of what people call "tech" sits in other NAICS codes — custom programming for a single client (541511), cloud hosting and data processing (518210), and ad-supported internet content (516210) — so the ~$492.9 billion is not "all U.S. software." In-house software that banks, retailers, and manufacturers write for their own use never shows up as industry output at all, and solo developers and open-source projects are largely absent [3].

Treat $592.3 billion as the size of the employer, primary-activity core of publishing — accurate as far as it goes, but an undercount of both content and software as economic activities.


4. Investable universe — where value concentrates across the children

(Tickers appear here and in Section 10 only.) The investable picture is where the two children diverge most sharply.

Software (5132) — deep, liquid, and unusually public. An outsized share of the value in this child is publicly traded and easy to reach. The child primer lists the major U.S.-listed publishers — Microsoft (Nasdaq: MSFT), Oracle (NYSE: ORCL), Salesforce (NYSE: CRM), Adobe (Nasdaq: ADBE), Intuit (Nasdaq: INTU), ServiceNow (NYSE: NOW), and dozens of pure-play cloud, security, data, and gaming names — plus diversified sector and cloud/cyber ETFs. The private layer is also huge: perennially private operators (Epic Systems, Bloomberg L.P., SAS Institute), venture-backed leaders (Databricks, OpenAI, Anthropic, Stripe), and the software-focused PE firms (Thoma Bravo, Vista Equity Partners, Silver Lake, Francisco Partners) that own scores of mid-sized platforms outright [3]. Because software is ~83% of the subsector, this child is where nearly all of 513's public-market investability lives.

Traditional publishing (5131) — scarce public pure-plays, clustered in two spots. Clean listed exposure is thin and concentrates in newspapers and scholarly-information houses: The New York Times Company (NYSE: NYT), News Corp (Nasdaq: NWSA/NWS), and USA TODAY Co. (NYSE: TDAY) among newspapers; RELX (NYSE: RELX), John Wiley & Sons (NYSE: WLY), and Informa among the scholarly/professional houses. Everywhere else — trade books (Penguin Random House, Simon & Schuster, Hachette), consumer magazines (Hearst, Condé Nast), data compilers, and the greeting-card duopoly — the durable value is private or foreign, reachable only as a segment inside a diversified owner or through private deals [2].

Bottom line for a public investor: the tradable, understandable exposure in 513 is overwhelmingly the software names (individually or via ETFs), plus a short list of newspaper and scholarly-information pure-plays in the content child. For a private investor the content child opens up considerably — scholarly and trade titles, backlist and rights libraries, proprietary data assets, and yearbook/card franchises are almost all private — while software adds a vast VC, growth-equity, and buyout opportunity set [2][3].


5. How the money works

Both children run on the same core: own the IP, license the access, reproduce at near-zero marginal cost. Once a title, database, backlist, or piece of code exists, incremental sales are enormously profitable, so scale, retention, and rights ownership decide returns in both worlds. And both have made the same structural shift — away from one-time, transactional revenue (a print ad page, a boxed software license) toward recurring revenue that repeats on its own (a subscription, a site license, a SaaS renewal).

The level of those economics is where the children part ways:

  • Software sits at the extreme. Gross margins often run 75–85%+, and the shift to recurring SaaS subscriptions makes revenue sticky and cash-generative. Investors judge these businesses on run-rate recurring revenue (annual/monthly recurring revenue, ARR/MRR), net revenue retention (how much existing customers spend this year versus last), the "Rule of 40" (growth rate plus profit margin above 40%), and free-cash-flow (FCF) margin — with AI-native products judged on margin after cloud-compute cost [3].
  • Traditional publishing's economics are far more varied, and its best pockets rival software: scholarly journals earn operating margins near 40% off unpaid authors and must-have library subscriptions; book backlist is a high-margin annuity; proprietary data compilation "looks like software"; and greeting cards carry very high product margins. Its weakest pockets — commodity print advertising, mail-order catalogs, paper maps — are postage- and paper-heavy and structurally shrinking [2].

The shared levers that decide whether high reported margins convert to cash: scale (to absorb fixed creation costs), the recurring-versus-transactional mix, and control of the cost base (paper and postage for content, cloud-compute for AI-heavy software).


6. Demand drivers

The two children answer to different customers, but the drivers rhyme:

  • Willingness to pay for access to IP. In both children, distinctive, must-have content commands direct payment while commodity content does not — a scholarly database, breakout fiction, verified data, or mission-critical enterprise software gets paid for; easily substituted output does not [2][3].
  • Digital transformation and automation drive the software child's secular growth: the multi-decade migration from on-premises systems to the cloud, plus rising corporate spending on data, analytics, and cybersecurity [3].
  • Institutional and R&D budgets drive the least-cyclical demand in the content child — library and university spending on journals, school budgets for textbooks and yearbooks, corporate budgets for business data [2].
  • The advertising cycle still moves the ad-dependent content pockets (local newspapers, consumer magazines) with GDP (gross domestic product) and consumer spending — but from a structurally shrinking base as budgets migrate to platforms [2].
  • AI — the subsector-wide swing factor. For software, AI is the biggest single driver, both a new product and a source of usage-based consumption; for traditional publishing, AI is an opportunity (licensing archives, building research tools) and a threat (zero-click answers that divert traffic). It is the one demand force that cuts across all of 513 [3][2].

7. Regulation

Publishing in both forms is lightly regulated relative to media and utility peers — no federal license to publish, no content mandate, no price control — and the rules that actually bite are cross-cutting and commercial rather than editorial. The active fronts differ by child but overlap on the biggest issues:

  • Copyright and AI — the shared central story. Copyright is what makes an archive, backlist, database, or codebase an asset, and it has moved to the center of the AI debate for both children. Landmark actions span the subsector: The New York Times Co. v. OpenAI/Microsoft and Bartz v. Anthropic (2025) — which held that training on lawfully acquired books can be fair use while pirated copies are not, alongside a ~$1.5 billion settlement, the largest U.S. copyright settlement on record — bear directly on content publishers, while software publishers face parallel disputes over training data, code, and model output [8][2].
  • Antitrust and merger review. The Department of Justice (DOJ) and Federal Trade Commission (FTC) review deals under the 2023 Merger Guidelines [7]. Enforcement is real in both children despite modest headline concentration: a federal court blocked Penguin Random House's purchase of Simon & Schuster in 2022 (content), and platform/software antitrust is at a decades-high (software) [2][3].
  • Data, privacy, and consumer-marketing law. State privacy laws (California's CCPA), the EU's General Data Protection Regulation (GDPR), FTC advertising and auto-renewal rules, and data-broker deletion regimes touch both children — sharpest on directory/data publishers in content and on data-handling software platforms [2][3].
  • Segment-specific rules. Postal (USPS) rate decisions are effective regulation for everything that moves by mail in the content child; export controls on encryption and advanced computing, app-store rules, and emerging AI regulation fall on the software child [2][3].

Net: regulatory risk is low relative to media and utility peers, but two levers — copyright/AI outcomes and antitrust — can reprice whole segments across both children.


8. Consolidation

By the subsector statistics this looks like an open field (CR4 20.2% [1]) — but, as Section 3 warned, that is a pooling artifact, and the real consolidation stories live inside the children, running in the same direction (buyers rolling up cash-generative assets) for different reasons:

  • Software consolidates on two engines: strategic buyers acquiring for reach and product (Broadcom's ~$69 billion VMware purchase; Salesforce's ~$27 billion Slack deal) and financial buyers — the software-focused PE firms — taking companies private and running them for cash. Take-private activity re-accelerated through 2025 after a 2022–23 lull [3].
  • Traditional publishing consolidates around its defensible pockets: newspapers merge among survivors and close the rest; a stable scholarly-journal oligopoly persists; the book "Big Five" is capped by antitrust while Amazon dominates distribution; and PE roll-ups gather data compilers, yearbook houses, and the greeting-card duopoly [2].

The common thread across all of 513: private equity and strategic buyers are consolidating the recurring-revenue, rights-defended assets, while commodity businesses are cut, sold, or closed. The subsector index hides all of it.


9. Risks

The two children share a risk spine, weighted very differently:

  • AI as a double-edged force — for software, AI can cannibalize per-seat pricing and raise compute costs even as it creates new products; for content, AI answer engines and AI-generated content can divert traffic and undercut undifferentiated work. This is the one risk common to the entire subsector [3][2][8].
  • Secular print decline — the existential risk in the content child (newspapers, consumer magazines, catalogs, maps); ongoing, not one-time [2].
  • Valuation and interest-rate sensitivity — the defining financial risk in the software child, where growth names are priced on far-future cash flows and re-rate sharply with rates [3].
  • Platform and channel dependence — reliance on a few gatekeepers: cloud providers, operating systems, and app stores for software; Google, Meta, and Amazon for content's traffic and distribution [3][2].
  • Copyright/AI litigation uncertainty — unresolved law on training and output cuts both ways for every IP-rich business in 513 [8][2].
  • Leverage at PE- and sponsor-owned assets — buyout debt loaded onto both mature content cash flows and mid-sized software platforms [2][3].
  • Measurement opacity — employer-only federal data undercount both children (nonemployer tail in content; reclassified output in software), and many of the most important operators are private, so top-down sizing is imprecise [2][3].

10. How to invest, and the outlook

513 is not a single trade — pick the child first, then the vehicle. The two children demand opposite playbooks.

For most public investors, 513 is a software story. Because software is ~83% of the subsector and unusually public, the deepest, most liquid menu is there: cash-generative mature names (Microsoft, Oracle, Adobe, Intuit) and higher-growth names (ServiceNow, CrowdStrike, Snowflake, Palantir) individually, or software and cloud/cyber ETFs for diversified exposure [3]. The content child adds only a short public list — newspaper pure-plays (NYT the cleanest scaled paid-digital winner; News Corp for premium business news) and scholarly-information houses (RELX, Wiley, Informa) as the most defensible recurring cash flows in traditional publishing [2]. Everywhere else in content, listed exposure is a segment inside a diversified owner, so the discipline is to isolate the publishing segment's revenue and operating profit from the parent before judging it.

For private investors, the opportunity set is far richer in both children. Software offers a vast venture, growth-equity, and buyout landscape (most software publishers are private); traditional publishing offers scholarly and trade titles, backlist and rights libraries, proprietary data assets, and yearbook/card franchises that sit almost entirely off the public market. The shared diligence weights recurring sell-through over catalog size, rights and clean chain-of-title over raw audience, and cash generation over accounting profit — and, for software, reconciling ARR against signed contracts and cash [3][2].

Outlook (judgment). Expect the two children to keep moving apart. Software should stay structurally attractive — extraordinary margins, sticky recurring revenue, demand that rises with automation — with AI widening the gap between firms that own durable, hard-to-replace workflows and those selling easily copied features. Traditional publishing should keep bifurcating internally — commodity print eroding while its recurring pockets (scholarly journals, book backlist, proprietary data, the card duopoly) hold up and attract private capital. The shared swing factors to watch are the net effect of AI (new licensing and product revenue versus cannibalized traffic and pricing), copyright and antitrust outcomes, and the ordinary technology and advertising cycles. No official industry forecast is provided in the federal statistics; 513 is best underwritten child by child — never off the deceptively blended 3-digit average. These are judgments, not guarantees.

For the full treatment of either child — complete company tables, unit economics, detailed regulation, consolidation, and per-metric sourcing — read the two child primers: 5131 Newspaper, Periodical, Book & Directory Publishers and 5132 Software Publishers.


Sources

This rollup is synthesized from the two child primers and our ground-truth federal statistics for NAICS 513. Per-metric citations for the children live in their own primers.

  1. U.S. Census Bureau, 2022 Economic Census — Concentration by Largest Firms, NAICS 513 — our ground-truth file for this level (receipts $592,340,719 thousand; 30,116 firms; CR4 20.2%, CR8 26.7%, CR20 35.6%, CR50 48.2%; HHI suppressed). https://www.census.gov/programs-surveys/economic-census.html
  2. Histometrics child primer, NAICS 5131 Newspaper, Periodical, Book, and Directory Publishers (receipts ~$99.4B; 13,326 firms; CR4 14.2%; HHI 110.2; five-child breakdown, investable universe, economics, regulation, consolidation, risks), synthesizing U.S. Census Economic Census 2022 and CBP 2023, industry association data, and company filings.
  3. Histometrics child primer, NAICS 5132 Software Publishers (receipts ~$492.9B; 16,824 firms; CR4 24.3%, CR8 32.1%, CR20 42.0%, CR50 55.3%; HHI suppressed; investable universe, SaaS unit economics, regulation, consolidation, risks), synthesizing U.S. Census Economic Census 2022 and CBP 2023 (child-level employment ~1.01M; payroll ~$203.7B) and company filings.
  4. U.S. Census Bureau, 2022 Economic Census — Concentration by Largest Firms (underlying child-level receipts, firm counts, CR ratios, HHI). https://data.census.gov/table/ECNSIZE2022.EC2200SIZECONCEN
  5. U.S. Census Bureau, County Business Patterns 2023 (child-level establishments, employment, annual payroll; employer-only universe; assembled to the ~1.26M-worker / ~$224B-payroll level figure cited here). https://www.census.gov/programs-surveys/cbp.html
  6. U.S. Census Bureau, 2022 NAICS Definitions — subsector 513 and industry groups 5131, 5132 (scope, exclusions, and the note that software publishing was numbered 511210 in the 2017 classification). https://www.census.gov/naics/?year=2022
  7. U.S. Department of Justice and Federal Trade Commission, Merger Guidelines, 2023 (HHI thresholds). https://www.ftc.gov/system/files/ftc_gov/pdf/2023_merger_guidelines_final_12.18.2023.pdf
  8. NPR / The Authors Guild, Anthropic settles with authors in first-of-its-kind AI copyright case; Bartz v. Anthropic (2025); and The New York Times Co. v. OpenAI/Microsoft — AI-and-copyright litigation across the subsector. https://www.npr.org/2025/09/05/nx-s1-5529404/anthropic-settlement-authors-copyright-ai