Skip to main content
search
AIData ManagementPodcastPodcasts

Useful AI Output Requires Quality Data Input: Stay Sharp Episode 45

By August 16, 2024December 17th, 2025No Comments

The 45th episode of Razorleaf’s Stay Sharp podcast, Data Quality and AI, features co-hosts Jen Ferello and Jonathan Scott digging further into artificial intelligence (AI) and its role in the manufacturing industry’s digital transformation. After starting the AI discussion in Episode 43, Under the Covers with AI, this week Jen and Jonathan focus specifically on the scope and the quality of the data your AI engine needs. In addition to discussing all the implications of the term “data quality,” they also address how companies need to maintain, test, and validate their data to keep AI generating good results.

Jonathan and Jen start with the caveats that they are not AI experts and they’re not trying to tell people exactly what they must do with AI. “I don’t think we mean to be prescriptive by what we’re talking about today, so much as informative. Think about these issues,” Jonathan says. “Do what you can to get ready.” Jen agrees, noting that the valuable part of discussions and consideration of AI at this early stage is to “recognize the importance at every step of the way of having good quality data and building an understanding of that.” Through their discussion, the duo also comes to realize that the topic isn’t only the quality of your data, but also the quality of your understanding of that data.

Data is the Foundation of AI

The reason it’s so important to really think about the data you are or will be feeding AI is that having the right data is really the first step, Jonathan says. “And if you get the first step wrong, you’re probably on the wrong path already.” AI must collect data with which it can recognize patterns and make decisions. “If it’s bad data, it’s the old adage, ‘garbage in, garbage out,’” Jen explains. “It’s got to be good quality data. All you’re going to be doing is making bad decisions really quickly.”

The duo agrees that “data quality” can mean a lot of different things. But they agree it goes way beyond the meaning of being complete and error free. Relevance of the data is one factor, Jonathan proposes, giving an example of a company that formerly had a more diverse product line but now makes only aircraft. “Would it make sense to include the nut and bolt design data from 30 years ago? I don’t know,” he says. Reliability of data is another important factor, Jen says. Is it predictable and reoccurring? How reliable was your input? Jonathan adds, “Can you trust it? Have you done some processing or steps in the middle that mean the data is not good or it’s only partly good?”

Representing the Real World

Jonathan suggests that when thinking about data quality, the key is asking yourself how accurately you’re representing the real world. Thinking about the context of what information you know to be true and where you’re making shortcuts can be critical. Because data scope is always going to be a consideration. “The world of possibilities—what you could do—is much wider than the world of practical capabilities—what you need to do,” Jonathan says. But you have to describe what to use and what to discard. Jen adds, “Just because you don’t need it now doesn’t mean that we won’t progress to the point that it does have value.” AI is extremely good at finding patterns and making connections that humans don’t see, and one of the potential pitfalls with AI is ensuring we’re not preventing those connections and patterns from being found because of the data we’re choosing to give it.

The Care and Feeding of AI

Jen and Jonathan also make the point that the relationship between your data and AI isn’t “set it and forget it.” AI requires testing, validation, and maintenance to continually ensure the results generated are still accurate, precise, stable, scalable, and reproducible. Jen says, “Because again, that foundational level, if the data is bad at any point, all you’re doing with AI is making bad decisions faster.” And since data changes over time, it’s important that you’re regularly maintaining and updating data, so that you keep getting good results and so that users continue to trust AI’s output.

Learn More About Data for AI

The full podcast episode offers many more details including how AI could be useful for your organizations regulatory compliance processes, considerations for balancing a model that meets all user needs with the complexity that data structure requires, the one must-do when working with data today, and more.

Check out the full conversation in Stay Sharp Episode 45: Data Quality and AI, and join us each week for a new podcast.

Follow the Razorleaf Podcast, Stay Sharp in Digital Engineering on:

Close Menu