FOXNUT

Definition

Spend analysis and the state of your own purchase data

What spend analysis actually runs on, why card transactions resist item-level classification, what the one public taxonomy offers, and what an analysis is worth once it must run without help.

By the Foxnut team · Updated

What spend analysis means here

Spend analysis is the work of taking every purchase a business has already made - purchase order lines, invoices, card charges, one-off payments - and classifying all of it into consistent categories, so that what is bought, from how many vendors, at what prices, can be answered from totals rather than from memory. It is the analytical half of what is sold as spend management: before anything renegotiates a contract or reroutes a payment, something has to say, reliably, what the money currently buys.

The search that leads here usually asks for tools, and the tools exist. What the shortlists do not say is that the deciding variable sits outside every product on them: the state of the purchase records the tool would run on. The same software produces a sharp analysis on clean records and a confident, wrong one on dirty records, which is why this page is about the records first and the automation second. It is also why the difference between a delivered analysis and an owned one matters more here than in most of this family - a spend cube is stale the quarter after it ships, so what an analysis system is worth once the consultant leaves is the standard to read before paying anyone to build one.

The demand under this topic is thin, and this page says so rather than working around it: most of what is typed into a search box here is either a request for a software shortlist, which this page declines to be, or another sense of the word analysis entirely. The question that remains is narrower and comes first - whether the purchase records a business already holds can carry any analysis at all.

The records the analysis runs on

Three populations of records feed a spend analysis, and they fail in three different ways.

RecordWhat it carriesWhat it does not carry
A purchase order or invoice lineVendor, amount, date, a description written for that one transaction, sometimes an internal account codeA category any other system shares, or descriptions consistent enough to group without interpretation
A payment card chargeVendor name as the processor renders it, amount, date, a merchant category codeAny item-level detail: the code classifies the seller, not the purchase
A one-off or unmatched paymentAn amount and a payeeA purchase order behind it, a vendor master entry, often any description at all

The card row deserves the most suspicion, because it looks classified and is not. A Merchant Category Code is, in the IRS’s own definition, a classification code assigned by a payment card organization to a merchant, “based on the predominant business activity of the merchant.” The code describes the seller’s main line of business. A charge at an office-supplies merchant may be furniture, coffee or software, and the record cannot say which.

The card channel is also not marginal spend. The United States government’s own charge card programme, GSA SmartPay, recorded $39.4 billion of spend in fiscal year 2025 at an average of $480 per transaction - the shape card purchasing takes everywhere, which is a very large count of very small purchases, each carrying merchant-level data only. The same statistics page records, with unusual candour, that the purchase card data a federal rule requires to be incorporated into the government’s procurement data system is currently not being reported into it, and is posted as spreadsheets on the programme’s website instead. Even the buyer with a statutory data standard has purchase records that do not reach the system built to hold them.

Classification is the actual work, and it has a public standard

Analysis understates the job. Before any total can be computed, every line has to be assigned a category, and the assignment is the hard part - the arithmetic afterwards is trivial. There is a public taxonomy for exactly this: the United Nations Standard Products and Services Code, a global classification of products and services that the UN’s own procurement marketplace uses to classify opportunities and suppliers. It is a tree of four levels - segment, family, class, commodity - two digits each, so a full commodity code runs to eight digits, and each step down the tree is one step more specific.

How well current AI classifies against that standard is measured, and the measurement is the most useful fact on this page. A 2024 study set GPT-4 to assign UNSPSC codes from item descriptions across a large, diverse dataset. Under its best prompt and settings, accuracy was 10.80% at the eight-digit commodity level, 29.01% at class, 40.31% at family and 54.59% at segment - roughly one correct code in ten at the bottom of the tree, rising steadily as the category coarsens. Earlier work with a fine-tuned classifier had reported excellent scores on one dataset and a fall to 50 to 55% when tested on a larger, more diverse one. And the study’s strongest prompt was the one instructed to say the information was insufficient rather than to guess, which is as much a statement about the input as about the prompt: a meaningful share of real purchase lines do not contain enough to classify.

The design consequence runs against the instinct to classify everything to the deepest level. A sourcing decision rarely needs an eight-digit code; it needs to know, dependably, that a family of spend is fragmented across nine vendors. Classifying to the level the decision needs, and letting the system abstain below it, is what the measured accuracy supports. Chasing commodity-level codes across the whole population is where these projects stall.

What the decision turns on

Six structural dimensions decide whether a process is worth automating. Spend analysis reads unusually on almost all of them, because it is the one process in this family whose input is the records every other process leaves behind, and whose output moves no money.

Spend analysis: process profile
DimensionWhat it reads on spend analysisSource
Exception varianceDefined by what a line fails to say, not by business events. The records that resist classification are whole structural classes rather than stragglers: card charges carry only a merchant category code, which the card organization assigns to the seller by its predominant business activity, so the line describes the merchant and never the purchase; invoice and purchase order lines carry free text written for one transaction; unmatched one-off payments may carry no description at all. The exception rate is therefore set by the mix of record types in the population, and it is knowable before any tool is bought, by looking.IRS Rev. Proc. 2004-43, sec. 3.03
VolumeWhole-population and batch, not a daily flow. A category total means nothing until every line in the period is classified, so the unit of work is the entire purchase history in scope, re-run whenever the taxonomy, the vendor master or the period changes. The card channel shows the shape of the count: the United States government's own charge card programme recorded $39.4 billion of spend in fiscal year 2025 at an average of $480 per transaction. The payback condition is fragmentation rather than any count: the work pays where third-party spend is spread across more vendors, systems and card statements than one person can read, and not before.GSA SmartPay program statistics, FY2025
Cost of an errorDiffuse, silent and downstream. No statute prices a misclassified spend line, no customer sees it, and no money moves because of it; a wrongly coded total still sums, which is why the failure makes no sound. The cost arrives through decisions built on the totals: a consolidation negotiated against a category that is really three categories, a supplier rationalised away that one site depends on. Measured accuracy makes the risk concrete: at roughly one correct commodity-level code in ten from an unreviewed model pass, most item-level codes in such a pass are wrong, and every total that includes them inherits the error without displaying it.Singh and Diao, IJCI 13(6), 2024
ReversibilityNear-total on the output, which is unusual in this family. A category code is an attribute in an analytical store: reassign it and re-run, and nothing outside the analysis needs to be told - no message was sent, no entry was posted, no filing window closes. That is why abstention is cheap here and guessing is not: a line marked insufficient waits for a person at no cost, while a guessed code propagates into totals, and the study's own best-performing prompt was the one that declared insufficiency rather than guessing. What does not reverse is a contract signed on wrong totals in the meantime: the analysis is correctable, and the sourcing decision made from it is not.Singh and Diao, IJCI 13(6), 2024
Regulatory exposureNone on the analysis itself, which is the sharpest contrast with this page's family. No rule requires a company to classify its own spend and nobody audits the result; the exposure that hangs over payments, postings and filings does not attach to a spend cube. The one regime that touches spend data at all governs disclosure of public money: United States law requires government-wide financial data standards - nonproprietary, searchable, machine-readable - and publication of federal spending on USASpending.gov. Even inside that regime the pipeline is incomplete: the government's own statistics page records that purchase card data required into the federal procurement data system at least annually is currently not reported there.Public Law 113-101, secs. 2 and 4; GSA SmartPay program statistics
Vendor market maturityA fixed public target with no settled answer at the bottom of it. The classification standard is mature: UNSPSC is a global taxonomy of products and services, and the UN's own procurement marketplace classifies suppliers and opportunities against it. The automation against that target is not settled: the published benchmark of LLM code assignment reads 10.80% at the eight-digit commodity level rising to 54.59% at segment, and the strongest earlier classifier reported mF1 of 0.912 on one dataset falling to 50 to 55% on a larger, more diverse one. A market with a stable standard and unsolved item-level assignment is one where the real differentiation is data preparation and review workflow, not the model.UNGM, UN Standard Products and Services Code; Singh and Diao, IJCI 13(6), 2024

Two rows set the design. The regulatory row says this is the rare money-flow process with no external referee - which removes the usual reason to keep a person in the loop and replaces it with a quieter and harder one: nothing outside the business will ever announce that the categories are wrong. Every other process in this family eventually gets corrected by a bank statement, an auditor, a counterparty or a regulator. A spend cube can be wrong for years, politely.

The reversibility and error rows form the working pair. The output is cheap to redo and the decisions made from it are not, so the review effort belongs at the level where decisions are made - the totals a negotiation will rest on - while the lines below are handled by classification with abstention, not classification with confidence.

What this usually gets wrong

The first error is buying the tool before looking at the records. The demo classifies the vendor’s sample data; the purchase runs on card statements where every line at a marketplace merchant is the same code, vendor names spelled four ways, and one-off payments with no description. The counts that decide the outcome are properties of the records, not of the product, and they can be taken in an afternoon without a vendor in the room.

The second error is classifying everything to the deepest level. Item-level assignment is the documented weak point of the automation and the least-needed output of the analysis; the decisions spend analysis feeds are made at family and segment level almost every time. Depth should be spent where a decision needs it, not distributed evenly across the population because the taxonomy has eight digits.

The third error is treating the analysis as a deliverable rather than a system. A classified cube handed over at the end of an engagement decays from the day it ships - vendors are added, systems change, the next quarter’s lines arrive unclassified. The durable asset is not the cube but the pipeline that produced it: the mapping rules, the vendor-name normalisation, the abstention queue and the person who clears it. An analysis that cannot be re-run by its owner is a photograph, and purchases keep moving after it is taken.

The fourth error is reading merchant-level data as item-level truth. A merchant category code answers who the seller mainly is, and a spend analysis needs to know what was bought; a report that quietly promotes one into the other will state, with precision, a breakdown of spend that nobody measured.

The verdict

The evidence supports automating the classification and the refresh, with abstention built in and review concentrated at the level where sourcing decisions are made. The measured accuracy says a model pass over item descriptions is a useful first sort at the coarse levels and unreliable at the bottom of the tree, which is an argument for a system that classifies to the level each decision needs, declares insufficiency instead of guessing, and re-runs cheaply as records arrive - and an argument against both extremes, the hand-built quarterly spreadsheet and the unreviewed automatic cube.

It does not support buying anything yet for a business that has not looked at its own records. The honest restatement of this page’s opening is that the shortlist most searches here want is the last step, and for records that fail the checks below it is not yet a step at all.

For a team weighing this, the first move is a count rather than a demonstration. Pull one month of purchase lines from every source - the purchase order system, the invoice ledger, the card statements - and count three things: lines whose only classification is a merchant category, vendors that appear under more than one name, and lines nobody in the room could classify without asking someone. Those three counts are the state of the purchase data. They decide whether the analysis can be automated, what it will cost to make the records carry it, and whether the tools question this page’s readers arrived asking is worth asking yet. Send us those three counts and we will tell you honestly which category your data is in.

Sources

  1. Internal Revenue Service, Revenue Procedure 2004-43 - section 3.03's definition of a Merchant Category Code as a classification code assigned by a payment card organization to a merchant/payee based on the predominant business activity of the merchant, and section 3.02's definition of a merchant/payee. Read as text extracted from the IRS's own PDF Retrieved
  2. United Nations Global Marketplace, 'UN Standard Products and Services Code (UNSPSC)' - the marketplace's own statement that UNSPSC is a global classification system of products and services, used by suppliers to classify what they sell and by UN staff to classify procurement opportunities, with its top-level segments listed A to J Retrieved
  3. Anmolika Singh and Yuhang Diao, 'Leveraging Large Language Models for Optimized Item Categorization using UNSPSC Taxonomy', International Journal on Cybernetics and Informatics 13(6), December 2024 - section 2.2's description of the UNSPSC tree (four levels, Segment, Family, Class and Commodity, two digits each, an eight-digit commodity code), tables 1 to 3 (GPT-4 accuracy under the best prompt at temperature 0: 10.80% commodity, 29.01% class, 40.31% family, 54.59% segment), the finding that the strongest prompt was the one that declared insufficient information rather than guessing, and the conclusion's account of earlier work reaching RoBERTa mF1 0.912 on one dataset and 50 to 55% on a larger, more diverse one. Read as text extracted from the arXiv PDF Retrieved
  4. Digital Accountability and Transparency Act of 2014, Public Law 113-101 - section 2's purposes, including Government-wide data standards for financial data and consistent, reliable, searchable spending data displayed on USASpending.gov, and section 4's data-standard requirements, including incorporation of widely accepted common data elements and a nonproprietary, searchable, platform-independent computer-readable format. Read in the US Government Publishing Office full text Retrieved
  5. US General Services Administration, GSA SmartPay 'Statistics and Reports' - fiscal year 2025 program statistics ($39.4 billion in total program spend, $471 million in refunds earned by agencies and organizations, $480 spent on average per transaction), and the page's own statement that FAR 4.606(a)(2) requires Government purchase card data to be provided at a minimum annually for incorporation into FPDS, that the data is currently not being reported into FPDS, and that it is posted on the GSA SmartPay website instead Retrieved

Foxnut Studios works on briefs like this one from Bengaluru and Paris. If you want the shape of that before you talk to anyone, here is how an AI engagement is scoped and priced.