Your AI Imported the CSV. Did It Keep the Customer IDs?

An automation can import every row, produce a clean summary, and still attach work to the wrong customer.

The quiet failure is often an identifier that looked like a number. A customer code such as 000417 becomes 417. A long reference number passes through a numeric representation that cannot preserve every digit. Two records that were distinct in the source become indistinguishable downstream.

Your agent may never see an error. The import succeeded. The spreadsheet opened. The CRM accepted the update.

For AI-assisted CSV imports, the practical rule is simple: treat identifiers as text, preserve their exact values, and check them before allowing customer-facing actions. A better prompt cannot reconstruct digits that disappeared before the model received the file.

An identifier is a label, not a quantity

Ask one question about each column: would arithmetic on this value mean anything?

Adding two invoice totals can be useful. Adding two invoice numbers is nonsense. Customer IDs, product SKUs, postal codes, order references, and tracking identifiers belong in the second category even when they contain only digits.

Consider this invented source file:

customer_id,customer_name,open_tickets
000417,North Studio,2
417,South Studio,1

Those customer IDs are different strings. Unless the source system explicitly defines them as equivalent, your workflow must not merge them.

If an import converts the identifier column to integers, both become 417. A later join might attach three tickets to one customer, duplicate rows, or overwrite an existing record. The exact consequence depends on the destination, but the original distinction is already gone.

This is a data-contract problem, not a model-intelligence problem. The tool-contract approach applies here: define what a field means before deciding how a tool should transform it.

Find the first conversion, not the last mistake

A typical workflow has more conversion points than its owner remembers:

  1. A source system exports a CSV.
  2. Someone opens and saves it in a spreadsheet.
  3. A script reads the saved file.
  4. An AI tool receives selected rows as JSON.
  5. A connector writes the result to a CRM.

Any step that guesses data types can change identifiers. CSV itself does not carry a universal column-type schema that every receiving application must obey. Quoting a value in the file is not a guarantee that a spreadsheet will retain it as text.

Keep the untouched export and compare the same known identifier at each boundary. If 000417 is intact in the export but missing zeros in the spreadsheet’s saved version, debugging the model is wasted effort.

Likewise, formatting a damaged cell as text afterward does not necessarily restore its original value. Once digits have been rounded or removed, retrieve the value from an authoritative source. Do not ask AI to invent a plausible repair.

Write a small import contract

You do not need a company-wide data platform. Start with the fields used to select a customer or trigger an action.

For an illustrative customer import, the contract could say:

  • customer_id is required text and must remain exactly as exported.
  • customer_name is display text, never the primary matching key.
  • open_tickets is a nonnegative integer.
  • Customer IDs must be unique within this particular customer table.
  • Unknown IDs go to review instead of creating or updating customers automatically.

Uniqueness belongs to the dataset’s meaning. A ticket table may legitimately contain the same customer ID many times. Do not reject valid repeated relationships because a generic validator assumes every column called “ID” is unique.

Also decide explicitly whether spaces and letter case matter. Automatically trimming whitespace or lowercasing identifiers is still a transformation. It is safe only when the source system’s rules establish that those changes preserve identity.

Keep the deterministic path away from AI

Use ordinary parsing and validation for the identifier fields. Let the model classify a note or draft a response, but carry the customer key alongside that work without asking the model to rewrite it.

Python’s standard CSV reader documentation describes the default behavior: fields are returned as strings, with no automatic type conversion unless specific options request it.

That makes a plain text-first reader a useful starting point. It does not make the whole pipeline safe. Later code can still call an integer conversion, a dataframe can infer a numeric column, or a destination field can reject text.

Pass the identifier as a JSON string, map it to a compatible destination field, and verify the stored value after writing. If your spreadsheet offers explicit import types, select text for identifier columns before conversion happens. Test the actual spreadsheet and connector combination your team uses.

Test the round trip with awkward examples

A file opening successfully is not an acceptance test. The useful question is whether its identifiers survive the complete route.

Create a tiny synthetic fixture containing:

  • A leading-zero ID such as 000417.
  • A distinct ID such as 417.
  • A long numeric-looking reference such as 123456789012345678.
  • An alphanumeric code such as AB-0042.
  • A deliberately empty required ID.

Send these through the same import, transformation, and export steps used in production, with external actions disabled. Compare the valid output identifiers to the input strings exactly. The empty required ID should be rejected, not filled with a guess.

Check row counts, unique-key counts where appropriate, and unmatched joins separately. Equal row counts cannot prove identity survived: both rows in the earlier example can remain while their keys collapse into one value.

If a check fails, stop before CRM updates, invoice delivery, or customer emails. Re-import from the untouched source after correcting the conversion boundary.

Make success mean preserved identity

A useful completion receipt says more than “CSV processed.” It states the input row count, accepted rows, rejected rows, unmatched identifiers, and whether the round-trip identity comparison passed.

The first implementation can be small: protect one export, declare one identifier column as text, and add five synthetic rows to a dry run. That is enough to expose a class of mistakes that polished AI summaries can hide.

Automation earns trust when it preserves the boring details. Before celebrating how quickly your agent handled the spreadsheet, make sure it still knows which customer is which.

More from the build log

Suggested

Want the full MarketMai stack?

Get the core MarketMai guides and operator playbooks in one premium bundle for $49.

View Bundle