Big Data Upgraded is the reading path for 23 articles, most of them first written in 2013. Each one is now rewritten for today’s tools. The original ran as 21 posts in October 2013. Most of its tool names have changed since then.

What Changed Since 2013
In 2013, big data meant Hadoop for most people. You stored files in HDFS and wrote MapReduce jobs or Hive queries. You moved data in with Sqoop and kept the cluster in line with ZooKeeper. Those names filled the original posts.
The picture looks different today. Cloud object storage replaced HDFS as the usual home for analytic data. Parquet files and open table formats such as Delta Lake and Apache Iceberg turned those files into real tables. Apache Spark took over the work people once wrote as MapReduce jobs. Sqoop was retired in 2021.
So I didn’t patch the old text. Posts on ideas that still matter were rewritten for the way teams work now. Posts built around one 2013 tool now teach the idea behind today’s platforms. Hadoop and Hive are still maintained, and Sqoop is retired. Most new platforms are built differently, so these names show up only as a line of history.
What Stayed the Same
The ideas held up better than the tools. Data still outgrows one machine. Work still gets split into pieces, moved by key and added up again. The relational database still sits at the center of most businesses, and it still feeds most analytics platforms.
That’s why MapReduce kept its place in the set, with a new diagram you can follow word by word. It’s also why the database posts are still here. NoSQL, NewSQL, document stores and graph models make more sense once you’ve seen what a relational engine does well.
Many posts now include a small T-SQL demo you can run on SQL Server 2025. Columnstore, the json data type, graph tables and change data capture all live inside one engine you already know. You don’t need a cluster to see most of these ideas work.
How to Read It
The map below groups the 23 posts into five stages. Start at the left and move right, or jump to the stage you need. Each stage builds on the one before it, but every post stands on its own.
Stage 1: The Foundations
These four posts give you the vocabulary. Read them first if the field is new to you.
- What Is Big Data? A Simple Explanation for Database People: the point where data outgrows one machine.
- The 5 Vs of Big Data: Volume, Velocity, Variety, Veracity and Value: five ways data gets hard, with a modern example of each.
- Structured, Semi-Structured and Unstructured Data Explained: tables, JSON and free text, and how each one is queried.
- Evolution of Big Data: From Flat Files to the Lakehouse: the history, and the problem behind each step.
Stage 2: Where the Data Lives
Storage changed more than anything else since 2013. These posts explain where analytic data sits today and why.
- Data Lake vs Data Warehouse: What Each One Is For: schema on read versus schema on write.
- Columnar File Formats: Why Parquet Changed Big Data: storing data by column, and why it saves so much.
- Open Table Formats: Delta Lake and Apache Iceberg Explained: how plain files became tables with transactions and history.
- Big Data Architecture Today: The Lakehouse and Medallion Layers: the whole platform, from sources to reports.
Stage 3: How Data Moves and Gets Processed
Once you know where data lives, the next question is how it gets there and how the work is split.
- Data Ingestion: Batch Loads, Change Data Capture and Streams: three ways data enters a platform.
- ETL vs ELT: How Big Data Changed the Way We Load Data: transform first, or load raw data and transform later.
- Batch vs Stream Processing: When Data Cannot Wait: scheduled jobs versus continuous streams, and the cost of real time.
- MapReduce Explained: Map, Shuffle and Reduce Step by Step: the pattern under every distributed engine.
- How Apache Spark Works: Driver, Executors and Partitions: the engine most teams use for big jobs today.
Stage 4: The Databases Behind It
This stage covers the operational side. It starts with the relational database and then walks through the other data models.
- Relational Database in Big Data: Still at the Center: why it still runs the business.
- OLTP vs OLAP: Why Transactions and Analytics Need Different Designs: two workloads, two designs.
- What Is NoSQL? Data Models, the CAP Theorem and Trade-Offs: the four models and the choice every distributed system makes.
- Key-Value and Document Databases: When Each One Fits: sessions, carts and profiles.
- Columnar, Graph and Spatial Databases: Three Shapes of Data: three more models, all inside SQL Server.
- What Is NewSQL? Distributed SQL Databases Explained: SQL and transactions across many machines.
Stage 5: Using the Data
Storing and moving data has one purpose: better decisions. The last stage is about getting answers you can trust.
- Big Data Analytics: Descriptive, Diagnostic, Predictive and Prescriptive: four kinds of questions.
- What a Data Scientist Does: The Skills That Still Matter: the work itself, step by step.
- Data Governance: Catalog, Lineage and Data Quality: knowing what you have and whether it’s right.
- Interview Question of the Week #022: How to Get Started with Big Data?: a short answer you can give in an interview.
Who This Path Is For
The posts assume you know SQL and have worked with a relational database. You don’t need any Hadoop or Spark experience. Every term gets explained the first time it shows up. Almost every post has a diagram that walks through one small example.
If you’re a DBA, stages 2 and 4 will feel closest to home. If you’re a developer moving toward data engineering, start with stage 3. If you manage a data team, read the architecture post and the governance post. Together they give you the whole picture.
What You Can Run Yourself
Twelve posts include a T-SQL demo that runs on SQL Server 2025. The free Developer edition is enough for all of them. Each demo creates its own small database and drops it at the end. Run them on a test server, never on production.
The demos cover columnstore storage, the json data type and graph tables. Others show distances with geography, change data capture, window functions and data quality checks. The other posts teach ideas you can’t run on one server, such as Spark or a distributed SQL database. For those, Big Data Upgraded leans on the diagrams and on small worked examples you can follow on paper.
If You Only Have an Hour
Read four posts. Start with What Is Big Data for the core problem. Then read the architecture post for the whole platform in one picture. MapReduce shows how the work gets split, and the data lake post explains where the results land.
You could argue that nobody needs this list anymore. A managed cloud service hides the storage, the shuffle and the file formats. You click a few buttons, and the queries run. That’s true until a query runs slowly or costs too much. Then the person who knows what happens underneath fixes it in an afternoon.
What to Remember
Big Data Upgraded keeps the ideas and swaps the tools. Most new platforms moved past the 2013 tool list, but the ideas came along. Splitting work, moving data by key and storing it by column all still apply. So do picking the right data model and checking that the numbers are right.
If you read the original series years ago, these posts will feel familiar and new at the same time. If you’re starting today, follow the stages in order. You’ll come out with the vocabulary and with the reasons behind it.
Big data is not a list of tools, it is a set of ideas that outlast any one tool.
Published by Pinal Dave on SQLAuthority. More of my work at pinaldave.com.
Discover more from SQL Authority with Pinal Dave
Subscribe to get the latest posts sent to your email.






6 Comments. Leave new
Hi,
I follows your series of posts regarding the big data.
but nowhere there is a post regarding the pro and cons, how to implement it, the type of competencies to implement it / use it (DBA, programmer, analyst?);
also, what are the standards?
there is a bunch of tools, none of them appear to have a common simple language like SQL or MDX.
What’s the cost of using this compared to traditional structure data “only”? (so structure/transform sooner (traditional DW) vs later (big data))
How the big data can fit the self service BI need and trend? (most of the business analysts have problems using structure data, so I cant imagine what’s unstructured data can be for them….)
how to handle the data quality issues? its fun to gather facebook, twitter and other information, but what’s the purpose if there is no quality in the data? is it the responsibility of each developer consuming these data to do and redo the job? how to compare this to traditional ETL processes? (remember: bad data quality = bad decisions!!!)
Personally I’m a little lost on where to put the big data in an enterprise. My background is DW oriented, with huge data quality and integrity issues, so structure data is a must have.
but it was a long and great job!
continue like this :)
Personally i feel no data is a bad data. I would call the process of deriving value out of data as a science in itself (Data Science). We have done projects where in data crawling was done from various social websites and output was generated which was meaningful. Regarding hosting of Big Data in an enterprise, we went with setting up a private cloud powered by Openstack and installed hadoop framework.
Great article, i liked this article as it gives simple definition of very commonly used words like PIG, HIVE, HADOOP. I think big data in combination with IBM Watson is going to be very big thing in future.
Great article, i liked this article as it gives simple definition of very commonly used words like PIG, HIVE, HADOOP. I think big data in combination with IBM Watson is going to be very big thing in future.
hi
thank you 4 your Big Data-serial article
many useful information!
I must say a big thank you to you for the post. But i have a question, can you show an example of a working model that transform unstructured data to structured data?. Thank you