Beyond the BIG DATA hype
One of the biggest business stories of the past few years has been that of big data. Academics, organizations, governments, and journalists have all become tied up with big data. Journalists cannot stop addressing enough stories about it, and viewers cannot watch enough coverage of the same. They all seem to be convinced: Big data is the remedy for all our hardships and difficulties. While some may be doubtful of the potential of big data, most have set their critical thinking aside and have voluntarily pounced on the big data bandwagon.
Today even those who cannot distinguish between standard deviation and standard error are singing the praise of big data and the assurance that it holds. I must confess, I have not yet drank the Kool-Aid. This is not to state that I am in disapproval of big data. Quite the contrary. I believe that big data is a relative term, and is more of a moving target and that it has always been around. What has changed today is the precise marketing around big data by the likes of IBM, SAS, and others in the analytics world, which has created mass awareness of what data can do for organizations, governments, and other entities.
The slick marketing has done a huge service to data science. It has converted the questioners by millions. There was a time when only the geek-dominated domains of computer science, economics, engineering, and statistics did one find individuals fascinated about data and statistical analysis. However, the continued marketing campaign about big data analytics, which IBM and others have spread across the earth in publications of high reputation since 2012, has made big data a household name.
The downside of this intense embrace of big data is the delusion that big data did not exist before. This would be a mistaken inference. In fact, big data has always been around. Even with the smallest storage capacity, any large dataset that exceeded the storage capacity was big data. Given the extensive development of the capacity to hold and manipulate data in the past decade, one could see that what was big data yesterday is not big data today, and will certainly not be regarded big data tomorrow. Today, even laptops are being shipped with hard drives of 1TB and more. I remember the days when my father used to carry a flash drive with a storage capacity of 512MB. What is important to realize is that any data is big data if its storage, manipulation, and analysis are beyond the capacity of the resources available to an entity.
It is believed that the future of computing will be cloud-based where individuals and organizations will store and analyze data online. At the same time, software will be available as a service rather than a product. In such circumstances, most analytics will take place in clouds such that the analysts will send commands for execution to the cloud and will receive output locally. Therefore, the size will become trivial because running the same analysis on small or big data will be no different.
The most important and significant point is to realize that what we do with data is more important than its size, or frequency, or complexity. We will not be limited by our ability to store or analyze data, but by our inability to interpret the empirical findings and devise strategies as a result.
When it comes to big data, it’s not the size that matters, but how we use it.
