What Is Big Data? It’s data that has outgrown what one machine and one database can handle in a useful time. Size alone isn’t the test. The test is the point where a single server stops being enough.

Why There Is No Magic Number
People who ask what is big data usually want a number, such as a terabyte or a petabyte. That number keeps moving. A 1 GB hard drive once felt endless, and a phone now holds hundreds of times more. A single server today can hold many terabytes.
So size alone is a poor test. A well-designed SQL Server database handles tables with billions of rows. Data becomes big when an ordinary approach starts to fail, and the failure shows up in one of four ways.
The first is storage: the data no longer fits on one machine. The second is time: a scan that should take minutes takes a day. The third is speed: data arrives faster than one server can write it. The fourth is variety: the data is JSON, logs, text and images, not neat rows. Only the first three are about size. Variety is about shape, and it calls for flexible formats, not more machines.
Three of the five Vs name these problems: volume, velocity and variety. The 5 Vs of Big Data: Volume, Velocity, Variety, Veracity and Value covers them in full.
What Each Failure Looks Like
Picture a pharmacy chain. Ten years of sales records from every till no longer fit on one server, which is a storage problem. The nightly report over all of them takes a whole day, which is a time problem.
Now picture a bike rental app with 200,000 bikes, each sending a location every second. That’s 200,000 writes a second, a load that one ordinary server struggles to keep up with. Then the company adds web orders as JSON, reviews as free text and photos of damaged bikes. The variety grows next to the speed.
Scale Up or Scale Out
When one machine runs out, you have two choices. Scaling up means buying a bigger machine, with more memory, faster disks and more cores. It’s simple, and your code doesn’t change. But there’s a ceiling, set by cost and by what the hardware can do.
Scaling out means adding more ordinary machines, called nodes, and giving each one a slice of the data. The ceiling is far higher, because you can keep adding nodes. The price is that the work must now be split and combined, which is harder than it sounds.
Scaling out also raises a safety question. Any node can fail at any time. So these systems keep copies of each slice on more than one node. That costs extra storage, but the data survives a failure.
One Worked Example
The diagram below uses 40 TB of data. On the left, one server is made bigger and bigger until it hits its ceiling. On the right, the same 40 TB is cut into 8 equal slices of 5 TB, one slice per node. A coordinator sends the query to every slice. Each node works on its own 5 TB and returns a small partial answer. The coordinator gathers the partial answers into one result.
Now the arithmetic. Say one machine reads 1 GB of data per second. Scanning 40 TB, which is 40,000 GB, takes 40,000 seconds, about 11 hours. Eight nodes scan their 5 TB at the same time. That takes 5,000 seconds, about 83 minutes.
The speed comes from the split, not from faster hardware. Real systems fall short of a perfect eight times, because the coordinator and the network take time too. The split and combine idea is the heart of tools like Spark. MapReduce Explained: Map, Shuffle and Reduce Step by Step walks through it on a small example.
How Today’s Tools Hide the Split
You’ll rarely split data by hand today. Apache Spark cuts large files into partitions and schedules a task for each one. Distributed SQL engines and lakehouse platforms do the same behind a SQL interface.
Formats such as Parquet store the slices as columnar files, so each node reads only the columns a query needs. You write SQL, and the platform plays coordinator.
Is Big Data a Marketing Label?
A fair objection: big data is a marketing word, and a tuned SQL Server solves most problems people call big. There’s a lot of truth in that. Before scaling out, check the indexes, the query plans, partitioning and columnstore.
Partitioning and columnstore can push one server a long way. Partitioning lets you move old data out in one fast step. Columnstore compresses data and scans only the columns a query uses. Try both first, and measure again.
Scaling out adds real cost. You take on a coordinator, a network and shuffles of data between nodes. Before anyone answers what is big data for your team, ask for proof that the single machine failed. Then ask whether it was storage, time or speed.
What to Remember
The answer to what is big data comes down to one test: does one machine still do the job? When storage, time or speed break the single machine, you have it. Scale up first when you can, and scale out when you must. Variety is a separate problem that needs flexible formats.
Ask what failed on the single server, and get a number. If nobody can name the failure, start with the query plan before adding nodes.
Big data is not a number of terabytes, it is the point where one machine stops being enough.
Published by Pinal Dave on SQLAuthority. More of my work at pinaldave.com.
Discover more from SQL Authority with Pinal Dave
Subscribe to get the latest posts sent to your email.






13 Comments. Leave new
Hi Pinal,
Good morning,
Thanks & Happy valentines Day! too
“I love who loves SQL…..!!!!!”
Happy valentines Day! too
I want to learn sqlserver please provide material
My first hard drive was a 10 MB (not a typo – MB!) “hard card” that I dropped into the expansion slot of my IBM PC (5150) some time in 1985. I had gone from 1982 until 1985 using the dual 5 1/4 inch 360K floppy discs and never thought I would fill up that 10 MB drive!
Hi ,
I am having one Date issue .For Eg if Date format is ‘2010-10-10 22:10:00.000’this is valid date and insrt to my database table. But if Date format is ‘2010-10-10 25:10:00.000’which is invalid date. It check the no. of hours is greter than 24 than it add one day to date and substract 24 hrs from time stamp.
Please advise.
Thanks in advance.
What is the datatype of the column that stores date values?
Thanks Pinal,
I was searching web to know about BIG data and hadoop, finally i end-up with your blog , simple and clar …
-Subbu
Thanks Penal,
Big Data is large amount of the data which is difficult or impossible for traditional relational database.
– simple definition .
These words enough to clear my Big Question of Big Data
Hi Pinal, Microsoft provided SQL Connector for Apache Hadoop (Linux) for SQL Server 2008 R2. Do we have a similar connector for SQL Server 2012 too ? How does SQL Server 2012 connect to Apache Hadoop on Linux ?
Four key characteristics that define big data:
>> Volume
>> Velocity
>> Variety
>> Value
Nonsense. These describe the capabilities of “Big Data” platforms; not the data itself.
For example, let’s say you have a workload that does not include “Volume”. By not meeting this metric, would you rule out using a Big Data platform?
his 4 keys don’t define BigData.. if anything.. big data is data that is too large to be processed quickly… so Volume goes hand in hand with big data,,, but Variety and Value do not equate to BigData
Hi Pinal Dave,
Actually I have executed one relationship as one hdfs table to one sql server table using sqoop export.
In the Hadoop is there any option export the data from
one hdfs table data to multiple sql server tables relationships,
many hdfs tables to one single sql server tables,
many hdfs tables to many sql server tables.
If there is possible to do the operations please reply for my question?