How to Prepare Your Data Before an AI Project
Most AI projects don't fail on the model. They fail on the data feeding it. Here's a practical guide for leaders on how to prepare data for an AI project before you spend a dollar on machine learning.
Frequently asked questions
It varies with data quality and complexity, but in most mid-market projects data preparation takes weeks, not days, and often represents the majority of the total project timeline. Clean, centralized data shortens this significantly, while data scattered across legacy systems extends it. The safest approach is to run a short discovery phase first so you can estimate realistically before committing to a full build.
Not necessarily. You can start a focused AI project by extracting and preparing a specific dataset without a full warehouse in place. That said, if AI is going to be an ongoing capability rather than a one-off, investing in centralized, well-governed data infrastructure pays off quickly across future projects.
Yes, but expect to spend real effort cleaning it first, and be honest about gaps. Some missing data can be worked around, while other gaps make a given problem unsolvable until you start collecting the right information. A good first step is an assessment that tells you what is usable now versus what needs to be gathered.
There is no fixed number, because it depends on the problem, the number of outcomes, and how strong the patterns are. For many structured business problems, a few thousand clean, representative examples are enough to start, while rare-event prediction needs more examples of the rare case specifically. It is better to start with quality data and expand than to wait for a perfect, massive dataset.
Ideally someone who understands both the business context and the data itself, often a data engineer or analyst working closely with the operational team that generated the data. The business context matters because only someone who knows how a field is actually used can spot when it is misleading. Clear ownership prevents the common failure where everyone assumes someone else validated the data.
Written by
Insights from the Codonomy team on custom software, AI, automation, and digital growth for B2B companies.
LinkedInGot a project in mind?
We build digital products that work. Let's talk about yours.