25 August 2026 · Pedro Aldea

508 Tables, 3M+ Records and 600,000 References: Data Before AI

A product-categorization case shows why AI activation starts by structuring the data, narrowing the decision and measuring the whole workflow.

When a company says it needs AI for its catalog, the useful question is not which model to choose. It is whether the data is ready for a decision that can be repeated, reviewed and improved.

In a product-categorization case, we worked with 508 tables, more than 3 million records and 327 brands. The challenge was not adding an AI layer to a spreadsheet. It was turning heterogeneous sources into an operational base the team could use without starting over every week.

The same work included more than 600,000 references that needed a consistent classification. The verifiable result was a categorization workflow executed in approximately 27 minutes. These are two different lessons: first, the scale of the data to structure; then, the runtime of the prepared workflow. Keeping them separate prevents a number from being sold as magic.

The problem was not a lack of AI

Large catalogs usually fail for very specific reasons:

  • product names written with different conventions;
  • duplicated brands or inconsistent variants;
  • attributes spread across tables with no clean shared key;
  • references with no category or incompatible categories;
  • exceptions known only by one person on the team.

Automating directly on that material can produce fast results that are hard to defend. Speed does not compensate for a classification nobody can explain or correct.

The first step was therefore to describe the sources, identify the columns that actually distinguished a product and define what a valid category meant. AI came later, where language interpretation and classification proposals were useful; deterministic rules handled the repeatable checks.

The sequence you can actually operate

A useful activation separates four layers:

  1. Inventory: which tables, rows, brands and fields really exist.
  2. Normalization: which names, units and keys need a shared form.
  3. Decision: which rule or signal assigns a category and which cases need review.
  4. Control: how the result is recorded, who corrects an exception and how the workflow runs again.

This sequence turns a data project into an operation. It also makes the risk visible: an ownerless table, an ambiguous field or an exception with nowhere to go.

What “27 minutes” actually means

A runtime is useful only if it can be repeated. That is why we do not present it as a universal promise or as the output of an isolated model. In this case, the workflow processed the prepared load, applied the agreed rules and left cases requiring judgment in a reviewable output.

The important measure is not only how long it takes. The team should also be able to answer:

  • which source produced a classification;
  • which rule or signal was applied;
  • which exceptions remain open;
  • what needs to change to make the next run better.

When those answers exist, AI stops being a demo and becomes part of the operation.

Is your catalog ready?

Before building, ask for five answers:

  1. What is the unit being classified: reference, variant, family or brand?
  2. Which fields are required and which values are valid?
  3. What share of cases has a clear answer and which exceptions appear?
  4. Who validates a proposal and how is the correction recorded?
  5. Which metric will be reviewed afterwards: coverage, consistency, time or errors?

If one answer is missing, the problem is not a missing model yet. It is a missing operational decision. In why we talk about operations, not AI we explain the full hierarchy: eliminate, simplify, automate and only then apply AI. For the data work that comes first, standardize before you automate is the natural follow-up.

A 15-minute test

You do not need to clean the whole catalog before finding the next bottleneck. Pick ten representative references from one source and record, for each one:

  1. which field or rule lets you make today’s decision;
  2. who would review an exception;
  3. whether the step is a rule, a lookup or a business judgment;
  4. which metric you will observe on the next run;
  5. which condition would make you stop and fix the data first.

At the end you have a one-page table. If several rows cannot be completed without asking one specific person, the bottleneck is in the data or the decision, not the model. That evidence is enough to decide whether to standardize one source, assign an owner or scope a small pilot before adding more automation.

The conclusion

The advantage is not saying that a catalog is large. It is being able to work with it without losing the data origin, the business criterion or the ability to correct.

508 tables and 3M+ records describe the scale of the problem. 600,000 references in roughly 27 minutes describe what a prepared workflow can do. Between those figures sits the work that cannot be skipped: inventory, normalization, decision and control.

That is the order that lets AI earn its place.