Data engineering does not need to be finished before AI can ship. It needs to be scoped to the specific use case the AI will serve. That distinction is the difference between a ten-week path to production and a twelve-month architecture programme that stalls.
Most mid-market CPG leadership teams told their AI programme needs "data engineering work" hear that as "a data warehouse project". It usually is not. AI Navi has written elsewhere on why data engineering is the real AI bottleneck for UK mid-market CPG and how the Flight Risk Index, AI Navi’s six-dimension delivery risk score, surfaces it before deployment. This piece answers the question that comes next: once the data engineering layer is identified as the constraint, how much of it actually needs to be fixed before the first working AI can ship?
How much data engineering do you actually need before shipping AI?
Enough to serve one use case. Not the portfolio.
AI Navi’s operating definition is deliberately narrow. Data engineering is ready when the pipelines can reliably move the right data, in the right format, at the right frequency, to support the specific AI use case the business has agreed to deploy first.
Not every future use case. Not the full data architecture the business will eventually need. One use case, defined narrowly, scoped to the data flows it depends on.
That constraint is what turns an abstract data problem into a solvable engineering problem inside a ten-week window rather than a twelve-month roadmap.
Why do most CPG data engineering programmes get scoped too broadly?
Because the wrong question gets asked first.
The question most consultancies ask is "what does your data architecture need to look like". The answer to that question is always large and always multi-year. It leads to a data warehouse project, a governance framework, a master data programme, and a sequenced roadmap that puts the AI ambition at the end.
The question AI Navi asks instead is "what data does this specific AI use case need to run reliably". The answer to that question is always smaller than the answer to the first one. It leads to targeted pipeline work, one or two governance decisions, and a working AI in production inside ten weeks.
Both approaches produce data engineering work. Only one produces a working AI at the end of a quarter.
What are the three data engineering failure patterns that stall CPG AI programmes?
Three patterns account for most of the stalled programmes AI Navi encounters inside £40M to £250M UK food, drink, and FMCG businesses.
- The missing middle layer. There is no transformation layer between raw ERP data and the model. The data arrives in the format the ERP produces, not the format the model needs. Someone assumed those two things were the same.
- The single point of knowledge. One person built the data pipeline. That person has since moved on, changed roles, or is too busy to maintain it. Nobody else understands it well enough to extend it. When the AI use case needs a new data feed, the project stalls while everyone works out who owns it.
- The governance vacuum. Two teams have different definitions of the same metric. On-shelf availability means something different to the commercial director than it does to the supply chain lead. The model is trained on one definition. The business makes decisions based on the other. Results look wrong and confidence collapses.
None of these are insurmountable. All of them take longer to fix than people expect when they have not been spotted early. The full diagnostic on how these show up on the Flight Risk Index sits in AI Navi’s earlier analysis, linked above.
How do you scope data engineering to a single AI use case?
Four questions, answered in sequence.
- One. What is the specific commercial outcome the AI use case is being deployed to move? Case fill rate. Deduction recovery. Forecasting accuracy. Trade spend effectiveness. Named outcome, named metric, named owner.
- Two. What data flows does that outcome actually require? Not the ideal state. The minimum viable set. Usually three to five feeds, not fifteen.
- Three. Which of those feeds already work, which need remediation, and which need to be built? Assessed against the specific use case, not the general architecture ambition.
- Four. Which governance decisions are blocking the model? Usually one or two, not a full framework. A definition of on-time delivery. A source of truth for the SKU master. An owner for one specific data feed.
That scoping conversation takes less than a week when it is scoped correctly, and it produces a plan a delivery team can execute inside ten weeks. It is the work AI Navi runs in the opening phase of the AI FlightPath™ Sprint.
How Do You Measure Your Data Engineering Maturity Right Now?
The honest answer: most CPG businesses don't know where they stand until something breaks.
We built our Data Engineering Readiness Assessment to change that a five-dimension diagnostic that takes your commercial, supply chain, and technology leads about fifteen minutes to complete. It scores your readiness across the five areas we've found consistently predict whether AI will ship or stall:
The Five Dimensions
| Dimension | What We're Testing | Common Failure Mode |
|---|---|---|
| Data Quality | Accuracy, completeness, and consistency across key datasets | Deduction data held in PDF attachments, not structured fields |
| Pipeline Maturity | How reliably data moves from source systems to decision-ready formats | Overnight batch processes that break without alerts |
| Team Depth | Whether you have the internal resource to own and maintain data flows | One analyst who built the whole thing and is now the only person who understands it |
| System Integration | How well your ERP, WMS, CRM, and planning tools connect | ERP and demand planning tool synced monthly, manually |
| Governance | Whether there are clear ownership, definitions, and controls on data | No agreed definition of 'on-time delivery' across commercial and supply chain |
Each dimension scores 1–10. The aggregate gives you your Data Engineering Maturity Score and we benchmark it against sector peers so you can see where you sit relative to other mid-market CPG businesses at similar revenue stages.
What Does a Typical Score Look Like in UK Mid-Market CPG?
Most businesses that come to us thinking they're AI-ready score between 3.5 and 5.5 out of 10 on aggregate. The highest scores tend to be on data quality (people generally know whether their numbers are clean) and the lowest on system integration and governance.
Pipeline maturity is where the surprises come. We've worked with businesses that had genuinely capable data teams, but whose pipelines had never been stress-tested under the load that AI requires specifically, the move from monthly reporting cycles to near-real-time inference. That jump exposes fragility that nobody knew was there.
The businesses that score above 6.5 are typically those that have already run one failed AI project and fixed things afterwards. Which is a more expensive way to learn than running a diagnostic first.
Maturity Score Interpretation
| Score Range | What It Means | Recommended Path |
|---|---|---|
| 7.0–10.0 | Data-ready for AI deployment | AI FlightPath™ Sprint — first working AI in production |
| 5.0–6.9 | Foundation in place, specific gaps to close | AI FlightPath™ Sprint with data remediation track |
| 3.0–4.9 | Structural data work needed before AI will hold | AI FlightScale™ Retainer — data engineering + AI in parallel |
| Below 3.0 | Current state will stall any AI initiative | AI FlightCheck™ diagnostic first to map exact gaps |
What Are the Most Common Data Engineering Failures in CPG AI Programmes?
Three patterns come up repeatedly. Not occasionally almost every time.
The missing middle layer. There's no transformation layer between raw ERP data and the model. The data arrives in the format the ERP produces, not the format the model needs. Someone assumed those two things were the same.
The single point of knowledge. One person built the data pipeline. That person has since moved on, changed roles, or is simply too busy to maintain it. Nobody else understands it well enough to extend it. When the AI use case needs a new data feed, the project stalls while everyone figures out who owns it.
The governance vacuum. Two teams have different definitions of the same metric. 'On-shelf availability' means something different to the commercial director than it does to the supply chain lead. The model is trained on one definition. The business makes decisions based on the other. The results look wrong and confidence collapses.
None of these are insurmountable. All of them take longer to fix than people expect when they haven't been spotted early.
How Do You Fix Data Engineering Gaps Without a Six-Month Programme?
You don't fix everything. You fix the specific gaps that block your specific AI use case.
This is the part where most consultancies get it wrong. They produce a full data architecture recommendation, a multi-phase, multi-year roadmap when what the business actually needs is three targeted interventions that make demand forecasting AI viable in the next ten weeks.
Our approach inside the AI FlightPath™ Sprint is to scope the data engineering work to the AI use case, not the other way around. We identify the exact data flows that the model needs, assess whether they exist and in what condition, close the specific gaps required, and ship the AI in production within the ten-week window.
The broader data architecture work may still need doing. But it doesn't need to happen before you get your first working AI in production.
A £40M food brand we worked with had convinced themselves they needed a full data warehouse before they could start. They didn't. They needed four specific pipeline fixes and a governance decision about one key metric. Eight weeks later, they had recovered 60% of their previously unchallenged deductions using a working AI in production built on the same systems they already had.
What does the ten-week path from stalled programme to production AI look like?
The AI FlightPath™ Sprint compresses the sequence.
- Weeks one to two. Scoping conversation, use case definition, and data flow assessment against that use case only.
- Weeks three to six. Targeted pipeline remediation, governance decisions closed, model development against the specified data.
- Weeks seven to eight. Model validation and integration into the operational workflow that will actually use it.
- Weeks nine to ten. Production deployment, handover to the internal team, and the operating cadence that keeps it running.
At the end, the business has a working AI in production, running on data flows engineered to support it specifically. Broader data architecture work may still need doing. It does not need to happen first.
Frequently asked questions
Do we need a data warehouse before we can ship AI?
Not for the first use case, in most mid-market CPG cases. Targeted pipeline work against the specific data flows the model needs is usually sufficient to reach production. Broader architecture work can be sequenced afterwards.
How long should data engineering scoping take?
Less than a week when scoped to a single use case. Longer scoping usually indicates the conversation has drifted into general architecture rather than use case data readiness.
What if we have already started a data warehouse programme?
The Sprint still applies. AI Navi has worked with businesses running a warehouse programme in parallel and still delivered production AI on separate pipeline work inside ten weeks, without disrupting the warehouse track.
Where to start
The AI FlightCheck™ diagnostic, market-comparable at $5,000 to $10,000 for equivalent audits, produces the scored Flight Risk Index and identifies whether a Sprint or a Scale retainer is the appropriate delivery route. Details at ainavi.co.uk. Your AI pilot worked in the sandbox, the data engineers said it was ready, the board approved the budget, and then nothing shipped.
If that sounds familiar, you're not dealing with an AI problem. You're dealing with a data engineering problem and it's the most common reason AI programmes stall inside mid-market CPG businesses right now.
At AI Navi, we work embedded inside food, drink, and FMCG businesses, and we've built a data readiness diagnostic specifically because we kept seeing the same breakdown: not in the model, not in the use case, but in the pipes that were supposed to feed it.
