The analogy needs a warning before I lean on it. A nutrition label and a model card are not the same document. One is regulated, standardized, and legally required. The other mostly is not: there is no common format, and labs disclose very different things. They share one useful property. Both are written by the people who want you to trust the product, and both are most revealing in what they carefully include and what they quietly leave out.
The “mostly” matters now. Since 2 August 2025, the EU AI Act requires providers of new general-purpose AI models to publish a summary of their training data; models already on the market have until August 2027. That is the ingredients list. The card format around it remains voluntary, and there is still no common standard. The analogy has shifted: in the EU, one nutrient is now on the mandatory label, while the rest of the card is whatever the maker chooses to show.
Here is how to read one.
Serving size. A nutrition label’s numbers mean nothing without a serving size. A model card’s numbers mean nothing without a description of what the model is for and how it was tested: the intended use, the prompt format, the effort setting, the evaluation harness. A headline number with the serving size stripped away is advertising. Good cards state this up front. Meta’s Llama 3 8B Instruct card opens by naming its intended use cases: commercial and research use in English.
Calories. This is the part everyone reads: the aggregate benchmark scores, the “beats the previous version” lines. Calories measure energy; they say nothing about health. A single average tells you nothing about the tails, the languages, the domains, or the failure modes. It is the number to start from, never the number to stop at.
The nutrient breakdown. Here is where the real reading happens. A good card separates safety into categories rather than one “safety score”: bias, toxicity, harmful content, cybersecurity, biological risk, jailbreak resistance, each with its own method and threshold. Some cards do this; others do not. Gemma 3 27B-it has an Ethics and Safety section with named areas such as child safety, content safety, and representational harms, but it reports results as a short prose summary of improvement over earlier Gemma models, without per-category numbers. Noticing that difference is the reading. When a card collapses safety into one paragraph, the collapse is itself information.
The ingredients list. Every food label names what went in. A model card should describe the training data, what was filtered, and how the model was aligned. This is often the vaguest section, and in the EU it is the one part now required by law. A card that will not say what the model was trained on is asking you to trust the dish without the recipe.
The allergen warnings. This is the known-limitations section: the specific ways the maker admits the model still fails. Read it before the benchmarks. It is one of the few parts written against the maker’s own interest. Meta’s Llama 3 8B Instruct card puts it plainly: “Testing conducted to date has been in English, and has not covered, nor could it cover, all scenarios.” That single sentence is worth more than any aggregate score, because it tells you exactly where the maker’s confidence runs out. The Gemini 3.8 Flash card places its Intended Usage and Limitations after the evaluation scores, which is why reading the warnings first matters.
What is not on the label. A nutrition label is required to list certain nutrients, so an absence on it is itself regulated. A model card has no such requirement, outside the EU training-data summary. Ask what benchmark is missing, which risk category was never measured, what was evaluated only by the maker’s team with no third-party reproduction. The omissions are often the most honest part of the document.
Consistency also counts. DeepMind keeps one model-card format across its Gemini versions at a single index page, which makes it easier to compare a new model against its predecessors. That consistency is a form of disclosure: when the format is stable, a reader can see what changed from version to version instead of decoding a new layout each time.
The practical rules follow:
- Read the serving size before the score.
- Read the warnings before the benchmarks.
- Look for the disaggregated rows, and distrust the single aggregate.
- Compare the ingredients sections across cards to see who discloses more.
- For every number, ask who ran the test and who else reproduced it.
- Treat the omissions as data.
Model cards are worth keeping. They are far better than nothing, and they are how a serious lab tells you what it checked. The point is to read them with the skepticism you bring to a label written by the manufacturer. The label is true as far as it goes, and what it does not say is also true, in its own way.

The Campfire
No commentsNobody has pulled up a log by this one yet. Be the first to say what you make of it.
Held for the desk. It appears after a look.