Missing Data Is Not Zero: Fix Your AI Report Before It Invents a Slump
Your morning report says yesterday produced zero leads. The AI recommends changing the offer, rewriting the landing page, and cutting the campaign.
The lead integration actually timed out.
Somewhere between the failed request and the finished paragraph, an unknown value became a zero. The model then explained a business event that never happened. A more persuasive prompt would only make the fiction sound better.
A reliable AI report must distinguish a measured zero from missing, partial, stale, and inapplicable data. That distinction belongs in the reporting pipeline before the model writes a sentence.
Zero is a result, not a fallback
Suppose a small agency sends clients a daily summary of leads, sales, and advertising spend. A successful query covering the intended account and full reporting window returns no qualifying leads. Under that metric’s definition, zero may be the correct result.
Now suppose the same query never completes. The count is unknown. Returning an empty list from the error handler does not make it a successful query with no matches.
A third case is subtler: the query succeeds, but only part of the requested window is available. Five observed leads might be accurate for that subset, yet misleading as the day’s total.
These cases need different representations and different prose. “No leads recorded yesterday” and “Yesterday’s lead count is unavailable” should never be interchangeable outputs.
Give every metric a collection state
Keep the number and its supporting state together. A useful metric record includes the value, unit, reporting window, collection time, source, and completeness state.
For example, a failed lead query could produce:
{
"metric": "qualified_leads",
"value": null,
"state": "unavailable",
"window": "2026-10-04 America/Chicago",
"source": "crm",
"reason": "request_timeout"
}
This is an illustrative contract, not a universal schema. Your implementation should store precise start and end timestamps and define which boundary is inclusive. The important rule is that the failure does not manufacture a number.
Use a small vocabulary with explicit meanings:
- Complete: collection met the metric’s coverage requirements.
- Partial: some observations arrived, but coverage is incomplete.
- Unavailable: collection did not produce a usable result.
- Stale: the available result belongs to an older collection period or exceeds your freshness limit.
- Not applicable: the metric does not apply to this account or workflow.
A complete result can still be wrong because of a bad filter or definition. State tracking makes those assumptions inspectable; it does not replace validation.
Find the places that silently fill blanks
Inspect error handlers, spreadsheet formulas, database queries, and chart formatting. Look for defaults that replace absent values with zero before anyone checks why the value is absent.
There is a useful database example in the PostgreSQL aggregate documentation. It explains that, except for count, the listed aggregate functions return null when no rows are selected. In particular, sum of no rows returns null, not zero. The documentation also describes substituting zero with coalesce when necessary.
That substitution can be perfectly appropriate. A confirmed, complete query finding no qualifying transactions may legitimately represent zero revenue under your definition. But the expression itself cannot prove the upstream collection succeeded.
Do not solve this by banning zero defaults everywhere. Place them after the evidence check. Preserve the original collection state so a formatting convenience cannot erase the difference between “nothing happened” and “we could not observe it.”
Refuse comparisons that lack a denominator
Imagine Monday had 12 qualified leads and Tuesday’s count is unavailable. The correct output is not a 100% decline. There is no supported Tuesday value to compare.
Likewise, if advertising spend is known but conversions are unavailable, cost per conversion is unavailable. If conversions are confirmed zero, dividing by zero still does not produce a useful finite cost per conversion. Report the spend and zero conversions separately.
Make these decisions deterministically before summarization. Require compatible reporting windows, units, metric definitions, and adequate collection states for each calculation. The model should receive an approved comparison or an explicit reason that the comparison cannot be made.
For mixed-quality reports, continue with the sections that remain supported. Missing CRM data does not necessarily invalidate a separately verified advertising-spend total. Avoid both extremes: inventing the missing number and suppressing every useful fact because one source failed.
Carry the uncertainty into the final sentence
Structured input alone is not enough if the writing layer strips away every qualification. Define acceptable language for each state.
A complete zero can become: “The CRM recorded no qualified leads in yesterday’s reporting window.” An unavailable value becomes: “Yesterday’s qualified-lead count is unavailable because the CRM request failed.” A partial value becomes: “Five leads were observed, but collection is incomplete; this is not the daily total.”
Those sentences tell the operator what happened and what remains unknown. They also prevent an unsupported recommendation from sneaking in after an accurate caveat.
If the pipeline cannot establish whether leads fell, it should not recommend changing the campaign because leads fell. It can recommend restoring collection and checking the result. This is the reporting equivalent of marking degraded sources explicitly, with the uncertainty carried through the arithmetic as well as the prose.
Test absence before trusting the dashboard
Build a small fixture set before enabling unattended summaries. Include a complete zero, a complete positive value, a timeout, incomplete collection, stale data, and a metric that does not apply. Add a comparison where yesterday is known and today is unavailable.
For each fixture, inspect three outputs: the stored metric, any derived calculations, and the finished report. A test fails if an unavailable value becomes zero, an unsupported percentage appears, or the prose turns uncertainty into a business diagnosis.
Start with your most consequential daily metric. Trace one missing observation from collection to the final sentence. Fix the first place where its meaning disappears.
Your AI should explain the business you actually observed, not the business your default values invented.
More from the build log
Suggested
Want the full MarketMai stack?
Get the core MarketMai guides and operator playbooks in one premium bundle for $49.
View Bundle