Guide

How to run a spend analysis

What data you need, how spend classification actually works, what a good result looks like, and the traps that quietly ruin most projects.

What a spend analysis is (and isn't)

A spend analysis answers three questions: what did we buy, who did we buy it from, and what should we do differently. It is not a reporting exercise. A dashboard that shows 40% of spend sitting in "Miscellaneous" tells you nothing you can negotiate with.

The work that makes the answer useful happens before the charts: pulling the transactions together, resolving supplier names to real organisations, and mapping every line to a category structure your team actually uses.

Step 1: Gather the data

Start with accounts payable or general ledger transaction lines, because that is where all spend eventually lands. For each line you want, at minimum:

  • Supplier name as it appears in the source system
  • Transaction date so you can compare periods
  • Value and currency, plus tax treatment if it varies
  • Free-text description — the single most valuable field for classification
  • Entity, cost centre or business unit for slicing later

Useful but optional: purchase order lines, contract references, supplier master records, and any existing category codes (keep them as a comparison, not as the truth).

Two or three years of history is plenty. Chasing a perfect extract before starting is the most common way a spend analysis stalls — begin with what you can export this week.

Step 2: Cleanse and normalise suppliers

In most extracts the same supplier appears several times: with a trading name, a legal name, a typo, a branch suffix, and a duplicate record created by a different entity. Until those are resolved, every supplier-level number you produce is wrong.

Normalisation matches those variants to one canonical supplier, and parenting rolls subsidiaries up to their ultimate owner — which is usually where negotiating leverage suddenly appears. Our supplier name normalisation and grouping & parenting tools do this step, and company AI enrichment adds registration numbers, sectors and group structure.

Step 3: Choose a taxonomy

A taxonomy is the category tree you classify into. You have three realistic options:

  • A standard schema (UNSPSC, eCl@ss, ProClass, CPV) — quick to adopt, good for benchmarking, often too granular or a poor fit for how your categories are actually managed.
  • Your own structure — mirrors how your category managers are organised, so the output is immediately actionable.
  • A hybrid — your structure at the top two levels, mapped to a standard underneath for benchmarking.

Three or four levels is usually enough. If a category has no owner, it does not need to exist. See taxonomy design for how we build and refine these.

Step 4: Classify the spend

Classification maps every transaction line to a category. Done by hand, it is a spreadsheet exercise that takes weeks, ages badly, and has to be repeated at the next refresh.

AI classification reads the supplier, the description and the context of each line, assigns a category and attaches a confidence score. The important part is what happens to low-confidence lines: they should be surfaced for human review, not quietly forced into a category. Decisions made in review then feed back in, so the next refresh is more accurate than the last.

AutoClass does exactly this — classifying against any hierarchical taxonomy in hours or days, with confidence scoring and review built in.

Step 5: Analyse and act

With clean, categorised data, the questions worth asking are straightforward:

  • Where is spend fragmented across many suppliers in one category?
  • How much sits outside contract or off-catalogue?
  • What does the tail look like — the thousands of small suppliers below your reporting threshold?
  • Which contracts are renewing in the next six months, and at what value?
  • Are we paying the same invoice twice?

Each of those has a follow-on: spend analytics for the reporting layer, tail spend analysis, contract management for renewals and obligations, and duplicate payment checking for recovery.

What good looks like

  • The large majority of spend by value classified, with the rest explicitly flagged for review rather than dumped in "Other".
  • Supplier counts that fall sharply after normalisation — that drop is the duplicates you were double-counting.
  • Every category has a named owner.
  • A refresh takes hours, not another project.
  • Category managers can trace any number back to the underlying transactions.

Where spend analysis goes wrong

  • Waiting for perfect data. Start with the extract you have; gaps show up faster in the output than in a data-quality workshop.
  • Trusting existing GL codes. They reflect accounting treatment, not procurement categories.
  • Classifying before normalising. You end up with the same supplier in three categories.
  • A taxonomy nobody owns. If no one is accountable for a category, no action follows the insight.
  • One-off exercises. A spend analysis that is not refreshed is out of date within a quarter.
  • Forcing every line to a category. An honest "needs review" bucket is more useful than false precision.

Frequently asked questions

What is spend analysis?

Spend analysis is the process of collecting, cleansing, classifying and reviewing your organisation's purchasing data so you can see what you buy, who you buy it from and where value is leaking.

What data do I need to start a spend analysis?

At minimum you need transactional data (invoice or ledger lines) with a supplier name, a date, a value and a free-text description. Purchase order data, contract records and supplier master data make the analysis stronger but are not essential to begin.

How long does a spend analysis take?

With AI-assisted classification, a first categorised view of spend can be produced in hours or days rather than the weeks or months typical of manual mapping exercises.

What accuracy should I expect from spend classification?

A well-designed taxonomy with AI classification and human review of low-confidence lines typically classifies the large majority of spend value, with the remainder routed for review rather than silently guessed.

Want this done on your own data?

We can take a spend extract and return a categorised, supplier-clean view in days — so you can see the shape of the opportunity before committing to anything.