Shopping Cart
Total:

$0.00

Items:

0

Your cart is empty
Keep Shopping

Why is good data more important than big data when training AI? (volume vs quality)

The language models we use work as statistical models. We feed them with "examples" and they try to guess the "logic" behind them (formulas, patterns). So to be as accurate as possible, it is said that quality is more important than volume. Quality is not just about feeding it with articles from the best journalist or writer, one of the priorities is also tagging and structured/consistent data. For example, if we want to train it with information about music albums from different artists, the best way should be to build an Excel document, where we have the same fields for everyone (year, gender, artist) and avoid repeating information, so the model can also group/filter them and be more accurate. If training it with raw text/articles, a consistent language style should be used. Volume, after meeting the quality requirement, is also what makes a model "powerful", but also complex, which translates in more resources (storage, processing units) to run it. For that reason, today, people create fine-tuning models, which is a technique to reduce their volume by removing useless data from it, and some open-source models (such as deepseek) can be optimized to run on any device.
Alvaro Lozano, Denmark, software-related entrepreneur, AI Educator
It should be noted that, no matter how good the data is, thousands and thousands of examples will always be needed to train an AI, i.e., a fairly considerable volume. Even so, it all depends on what we consider quality and what quantity we are talking about. Obviously, junk data that does not correspond at all to what we want the AI to do is not going to help us. Likewise, between having 100,000 cases of decent data that broadly cover the cases we want the AI to cover, or having 1,000 perfect cases of the highest quality, 100,000 is better, because with 1,000 you achieve practically nothing. It is necessary to find a balance between quantity and quality so that the model can generalize correctly. In any case, good data is considered more important. This is because, as I said, AIs need examples to learn to do what we want them to do. If you want them to program well, you have to give them well-programmed code, and if you want them to draw giraffes, you have to give them giraffes and not elephants. In short, data sets the limit of what an AI can learn.
Eneko Garzon Uria, Spain, computer engineering student, AI expert
Because AI ultimately learns from data, the quality of that data has a direct impact on how well it performs. If the data is accurate, relevant, and well-organized, the AI is much more likely to produce useful and reliable results. However, having more data doesn’t automatically mean better outcomes. Large amounts of low-quality or biased data can actually confuse the system and lead to poor performance. What really matters is having data that is clean, meaningful, and representative of what the AI is trying to learn. In the end, it’s not just about how much data you have, but how good that data is.
Leticia Sagredo, Spain, trainer on professional retraining and the acquisition of digital skills
When training artificial intelligence, good data is often more important than the largest possible amount of data. It is often assumed that AI automatically gets better when it is trained with "big data". However, the quality of the data is crucial. If data is incorrect, incomplete, or one-sided, the AI learns incorrect patterns and may make inaccurate or unfair decisions. The AI then passes this on as an answer – and may therefore not be usable. Good data is characterized by being accurate and well-structured. A smaller, carefully selected amount of data can therefore be more valuable than millions of unverified pieces of information. This can be compared to learning at school: students benefit more from clear, understandable and correct learning materials than from a large amount of disordered content. For me as a teacher, this topic shows how important critical thinking is when dealing with AI. Students should understand that AI is not "objective" or "omniscient", but only works as well as the data with which it was trained. Bad data can reinforce biases or generate false results. That is why responsible handling of data, transparency and human control are needed. The quality of the information thus forms the basis for trustworthy and fair AI systems.
Susanne Fürstner, Austria, AI Enthusiast
Well imagine training all your life for ice skating in the olympics but finding out when you finally go there that the thing you were training all this time for was actually hockey. It doesn’t really matter how much training you give a model, when the data is flawed or incorrect.
Juan Agustin Veliz, Germany, AI Expert
In my opinion, high-quality data is more important than large amounts of data when training AI, because quality always trumps quantity. When there is a massive amount of information, it can often turn out to be inaccurate or incomplete, or even incorrect, and thus the model will learn from these errors. This can lead to inaccurate results and biased responses. A smaller, curated, structured dataset can provide AI with a much more organized and clear foundation for understanding patterns. A much more significant downside, of course, is the resources required to process the data. With larger and more complex data, much greater amounts of energy, money, time, and natural resources are needed. With a smaller dataset, this is significantly reduced, in my opinion. Additionally, the objectivity of the AI model is extremely important, and with a smaller amount of data, it can learn to be more objective and avoid the risk of biased responses. It is important to provide it not with all the information, but with the right information.
Nikolai Dimitrov, Bulgaria, AI Enthusiast
0
Would love your thoughts, please comment.x
()
x