# Prebiotic & Fiber Label Index — method Gut Health Times, built 2026-09-22. This file documents exactly how `fiber-index.csv`, `fiber-index-summary.json` and `fiber-index-summary.md` were produced. Nothing in them is estimated, modelled or filled in by hand. Every figure is computed by `build-fiber-index.py` from one public file set, and the script is the authority: where this file and the script disagree, the script is right. --- ## 1. The source | | | |---|---| | Dataset | USDA FoodData Central, **Branded Foods**, CSV release | | Release date | **2026-04-30** (the release USDA labels `2026-04-30`; the zip's own `Last-Modified` header is Wed, 29 Apr 2026 22:59:39 GMT) | | URL | | | Downloaded | 2026-09-22, ~18:43–18:45 UTC | | Size | 448,767,220 bytes zipped; 3,091,857,579 bytes unzipped across 11 files | | SHA-256 of the zip | `26050a5d03197469813754743a21ee0fad4ccf22b6aac2a995846a987719fc49` | | Files actually used | `food.csv`, `branded_food.csv`, `food_nutrient.csv`, `nutrient.csv` | | Listing page consulted | (2026-04-30 was the newest Branded Foods release offered) | The **full** dataset was downloaded and processed. Nothing was sampled, and the FoodData Central API was not used at all — no API key was requested and no account was created. The zip is not committed to this repo; re-download it with the command in section 8. Nutrient ids were read from the release's own `nutrient.csv`, not assumed: | id | name | unit | |---|---|---| | 1079 | Fiber, total dietary | G | | 2000 | Total Sugars | G | | 1063 | Sugars, Total | G (older label records use this id; used only when 2000 is absent) | | 1235 | Sugars, added | G | | 1008 | Energy | KCAL | | 1003 | Protein | G | | 1005 | Carbohydrate, by difference | G | --- ## 2. Filters, in the order the script applies them 1. **Market.** Keep `branded_food.market_country` in `{"United States", "US"}`. The release also carries 1,117 New Zealand records; they are dropped. Both US spellings are kept because the dataset uses both. 2. **Qualifying.** A label record qualifies if **either**: - **(a) a named fiber ingredient** — the product name (`food.description`) or the ingredient statement (`branded_food.ingredients`) matches one of the 14 ingredient patterns in section 3; **or** - **(b) the 2 g rule** — the record sits in a beverage / bar / cereal / snack / yogurt category (the exact category list is the `CAT_*` sets in the script) **and** declares **≥ 2 g** of dietary fiber per labelled serving. Rule (b) deliberately catches fiber-forward products that never name a fiber ingredient (whole-grain cereals, bran, popcorn, chia). Every row records which rule it qualified under in `qualified_by`, so the two populations can be separated at any time. 3. **De-duplication.** FoodData Central keeps historical snapshots of the same label: one UPC can appear 25+ times across releases, each with its own `fdc_id` and publication date. `fiber-index.csv` keeps **every** qualifying record (so label history is visible and auditable), and marks exactly one per `gtin_upc` with `is_latest_for_upc = yes` — the record with the latest `publication_date`, ties broken on the higher `fdc_id`. **Every number in the summary is computed on that de-duplicated set**, i.e. one row per UPC. Records with no UPC would each be kept, but in this release every qualifying record has one. 4. **Plausibility.** A record is excluded from every statistic (but kept in the CSV, with `included_in_stats = no` and the reason in `data_quality_flags`) if any of these is true of its per-100 g figures — each is physically impossible, not a judgement call: - `fiber > 100 g` — more fiber than the food weighs; - `fiber > carbohydrate + 1 g` — dietary fiber is a component of total carbohydrate; the 1 g tolerance absorbs label rounding; - `carbohydrate > 100 g`, `protein > 100 g` or `total sugars > 100 g`; - `carbohydrate + protein > 101 g` — more mass than the food has; - `energy > 900 kcal` — 100 g of pure fat is 884 kcal. Nearly all of these are records where the provider filed per-serving figures in the per-100 g fields. Without the gate the maximum "fiber per serving" in the release is 8,888 g — a cold-brew coffee whose per-100 g fiber is recorded as 2,469 g. **The gate removes the impossible, not the improbable.** Records that are internally consistent but still wrong survive it, and they cluster in the top of the distribution. The maximum and the "highest declared fiber" table are therefore unaudited: treat them as a list of records to check by hand, never as a finding. The median, the quartiles and the share thresholds are the robust numbers and are barely moved by the tail. No other exclusion is applied. Discontinued products are **kept** and flagged (`discontinued_date`), because a discontinued label is still evidence of what the category shipped. --- ## 3. Fiber-ingredient detection Patterns run against a lower-cased, whitespace-collapsed concatenation of the product name and the ingredient statement. A product matching several patterns is listed under all of them, so ingredient counts sum above the product count. | label in the CSV | matches | |---|---| | inulin | `inulin` | | chicory root | `chicory root` (not when preceded by "roasted"), `chicory fiber/inulin/extract` | | agave inulin | `agave inulin / fiber / fructan` | | tapioca fiber | `tapioca fiber`, `fiber (tapioca` | | soluble corn fiber | `soluble corn fiber`, `corn fiber syrup` | | IMO | `isomaltooligosaccharide`, `isomalto-oligosaccharide`, standalone `imo` | | resistant dextrin | `resistant [tapioca/malto/potato/corn/wheat] dextrin`, `fibersol`, `nutriose`, `promitor` | | polydextrose | `polydextrose` | | acacia | the bare word `acacia` (see the caveat below) | | psyllium | `psyllium` | | beta-glucan | `beta-glucan` / `beta glucan` | | cassava root fiber | `cassava root fiber`, `cassava fiber` | | citrus fiber | `citrus fiber` | | pea fiber | `pea fiber` (word boundary stops `chickpea fiber`) | Three decisions worth stating plainly, because they move the numbers: - **Plain `tapioca dextrin` is NOT counted as tapioca fiber.** In this release it appears 13,169 times, almost always inside the "contains 2% or less of" line beside calcium stearate, glycerin and flavours — it is a coating and carrier, not a declared dietary fiber. Only `tapioca fiber` (2,490 records) counts, plus `resistant tapioca dextrin`, which counts as a resistant dextrin. - **`roasted chicory root`** in coffee and coffee substitutes is a flavouring, not a fiber source, and is excluded by a negative lookbehind. - **Acacia is counted, but the form is recorded separately.** "Acacia" is the term the brief asked for, and it comes out as the single most-named ingredient in the index. But `gum acacia` (20,900 records) and `acacia gum` / `acacia (gum arabic)` (9,505) vastly outnumber `acacia fiber` (749). The gum form is a glaze, emulsifier and flavour carrier declared under the 2%-and-under line; its presence is not a fiber dose. Every acacia product therefore also carries an `acacia_form` value — `gum`, `fiber`, `fiber+gum` or `unspecified` — and the summary reports the split. **Do not publish the raw acacia rank without that split.** Marketing terms are a separate closed list (`prebiotic`, `probiotic`, `high fiber`, `excellent source of fiber`, `good source of fiber`, `fiber-rich`, `added/plus fiber`, `gut health`, `digestive health`, `microbiome`, `regularity`, and a bare mention of the word `fiber`), detected on the same text. They are recorded twice: `marketing_terms` (name + ingredient statement) and `marketing_terms_in_name` (product name only). The naming-vs-grams comparison uses the name-only field, because that is what a shopper reads. --- ## 4. Normalisation - **Per-serving values.** FoodData Central stores every branded nutrient **per 100 g**. `fiber_g_per_serving = fiber_g_per_100g / 100 × serving size`. - **Millilitres.** `serving_size_unit` of `ml` / `MLT` is treated as grams, 1:1. That is the dataset's own convention, and it checks out: OLIPOP's record is 2.5 g/100 g on a 355 ml serving → 8.88 g, against a 9 g printed panel. - **Unit spellings.** `g`, `GRM`, `GM` are grams; `ml`, `MLT` are millilitres. `MG`, `IU`, `MC` are supplement doses that cannot be converted — those records keep their per-100 g values and get a blank per-serving figure. - **Percentiles** use linear interpolation between closest ranks (numpy `linear`, R type 7). The formula is `quantile()` in the script. - **Rounding.** Everything is rounded to 2 dp at output only; all arithmetic runs at full float precision. Because per-serving values are *recomputed* from a rounded per-100 g figure, they will differ slightly from the printed panel — a panel that says 3 g may compute to 3.02 g here. Never present a computed per-serving value as the printed label figure for a single product. - **Brand owners are not rolled up.** `GENERAL MILLS SALES INC.` and `General Mills, Inc.` are separate rows, exactly as the dataset spells them. --- ## 5. Known limitations — read before publishing any of this 1. **The labels are self-reported and often stale.** Records are supplied by manufacturers and data providers (mostly Label Insight, `data_source = LI`). USDA does not audit them. The median qualifying product's label record was published in **May 2022**, and **62.4%** (34,180 of 54,753) were published before 2023. Several entries in the summary's own cross-check are three years out of date. 2. **It is not a census of the market.** It is a census of *what has been submitted to FoodData Central*. Fast-moving DTC brands are patchily covered and several 2025–26 launches are simply absent. Never write "there are N products on the US market" — write "N label records in USDA's branded catalogue". 3. **A product's presence and its current formula are different questions.** The Poppi records in this release carry the pre-2025 recipe. Verify any single product against the live panel before writing about it. 4. **Serving sizes vary and some are placeholders.** Some records carry a serving size of exactly 100 g where the provider had no label serving; that inflates every per-serving figure on that record. Those rows are flagged (`data_quality_flags` contains `serving_size=100g`) and kept, not removed. Cross-product per-serving comparisons are comparisons of *labelled servings*, which is the honest unit for a shopper but not a fixed-weight comparison — the per-100 g columns are there for that. 5. **Ingredient text is free text.** A brand that writes "prebiotic fiber blend" without naming the fiber will not be detected. Detection is a floor, not a ceiling. 6. **Ingredient presence is not dose.** The index records which fibers a label names and how much total fiber the product declares. It cannot say how much of that total came from any single ingredient — no US label discloses that. The per-source medians in the summary are the *whole product's* fiber among products naming that ingredient. 7. **Added sugars are missing from most records.** Only 20,222 of the 54,753 qualifying products (37%) declare an added-sugars value, so the added-sugar bands are computed on a self-selected subset (broadly, records refreshed after the current Nutrition Facts panel became mandatory). 8. **No compliance claims.** The summary compares names against the FDA thresholds in 21 CFR 101.54 (a "high fiber" claim needs ≥ 20% DV, i.e. ≥ 5.6 g on the 28 g Daily Value; "good source" is 10–19% DV). This is **not** a compliance finding: FDC gives a product name and an ingredient statement, not the principal display panel, and the per-serving value is recomputed rather than read off the panel. Report it as a gap between naming and grams, never as a violation. 9. **The tail is not publishable.** See section 2.4. Percentiles up to the 99th are stable; the 99.9th percentile and the maximum are not. 10. **Categories are the dataset's own.** `branded_food_category` is assigned by the data provider and is sometimes wrong (a coffee elixir filed under "Soda"). The group mapping in the script is ours and is listed there in full. 11. **Brand-owner strings can be wrong at source.** Poppi records in this release are attributed to a brand owner called "Sew Sensational", and OLIPOP records to "Deco Home Solutions" and "Sahara Beans LLC". Match on `brand_name` as well as `brand_owner`, and never present `brand_owner` as a corporate fact. --- ## 6. Cross-check against five hand-verified panels Five products Gut Health Times had already verified by hand — from manufacturer or retailer panels, on the dates recorded in the script — were matched back against the release. None of those five figures comes from FDC. Matching is two-stage: every label record whose brand owner / brand name / product name / subbrand matches the brand pattern inside a plausible category, then those whose product name also matches the product pattern. Verdicts: **MATCH** = at least one matching record within 0.5 g per serving of the verified panel; **MISMATCH** = none is; **ABSENT** = not in this release. Results, with per-record evidence, are in `fiber-index-summary.md`, in the `hand_verified_cross_checks` block of `fiber-index-summary.json`, and row-by-row in `fiber-index-crosscheck.csv`. Outcome of the 2026-09-22 run — **two of five match, two mismatch, one is absent**, which is itself the strongest evidence for limitation 1: | product | GHT hand-verified | FDC in this release | verdict | |---|---|---|---| | Poppi, 12 fl oz | 3 g (cassava root fiber + agave inulin) | 18 records, all 2.16 g, agave inulin only, newest published 2023-10-26 | **MISMATCH** — FDC still carries the pre-2025 recipe | | OLIPOP, 12 fl oz | 9 g refrigerated | 61 records, 5.99–9.94 g, newest 2023-09-14 | **MATCH** on the 9 g line; the 6 g shelf-stable line and the 3 g MINIs are not in the release | | Trader Joe's Apple Prebiotic Soda | 5 g agave inulin | 141 Trader Joe's records in the release, none of them this product | **ABSENT** | | Halfday iced tea, 12 fl oz | 6 g (cassava root fiber, acacia, agave inulin) | 12 records, all 8.16 g, agave inulin only, newest 2023-11-16 | **MISMATCH** — FDC carries an older formula | | Optimum Nutrition Protein Shake + Fiber | 8 g soluble corn fiber | 95 Optimum Nutrition records, none naming fiber | **ABSENT** | Never source a single product's fiber figure from this index. Source the category from it, and the product from the live panel. --- ## 7. Outputs | file | what it is | |---|---| | `fiber-index.csv` | one row per qualifying **label record**; `is_latest_for_upc = yes` marks the one row per UPC that every summary number is computed on | | `fiber-index-summary.json` | every headline number, plus a `_methods` object giving the exact rule used for each key | | `fiber-index-summary.md` | the same numbers, readable, with the method printed under each | | `fiber-index-crosscheck.csv` | the raw matched records behind the five hand-verified spot checks | | `build-fiber-index.py` | the script that regenerates all four from the source zip | | `build.log` | stderr from the run that produced the current outputs | Column notes for `fiber-index.csv`: `serving_grams_used` is the serving size the per-serving columns were computed with (blank where the unit could not be converted); `fiber_ingredients` and `marketing_terms` are `; `-separated; `acacia_form` is populated only for acacia products; `included_in_stats` is the plausibility gate from section 2.4. --- ## 8. Reproduction ```sh curl -L -o branded.zip \ https://fdc.nal.usda.gov/fdc-datasets/FoodData_Central_branded_food_csv_2026-04-30.zip shasum -a 256 branded.zip # 26050a5d03197469813754743a21ee0fad4ccf22b6aac2a995846a987719fc49 unzip -q branded.zip python3 build-fiber-index.py \ --src FoodData_Central_branded_food_csv_2026-04-30 \ --out . ``` Python 3.9+, standard library only, no network access at run time. The run takes about 6–7 minutes and about 1 GB of RAM on an Apple-silicon Mac; it streams all three large CSVs rather than loading them. **For the next USDA release** (Branded Foods ships roughly every six months): change `RELEASE`, `DATASET` and `DATASET_URL` at the top of the script, point `--src` at the new folder, and re-run. Nothing else needs to change. Keep the previous `fiber-index.csv` — the release-over-release diff is the part of this asset nobody else can copy.