Big Data Upgraded: The 2013 Series Rewritten for Today

Big Data Upgraded is the reading path for 23 articles, most of them first written in 2013. Each one is now rewritten for today’s tools. The original ran as 21 posts in October 2013. Most of its tool names have changed since then.

Gouache painting of an old pencil sketch of a small building on worn paper beside a clean pale wood model of the same building, with a vermilion pencil resting between them.

What Changed Since 2013

In 2013, big data meant Hadoop for most people. You stored files in HDFS and wrote MapReduce jobs or Hive queries. You moved data in with Sqoop and kept the cluster in line with ZooKeeper. Those names filled the original posts.

The picture looks different today. Cloud object storage replaced HDFS as the usual home for analytic data. Parquet files and open table formats such as Delta Lake and Apache Iceberg turned those files into real tables. Apache Spark took over the work people once wrote as MapReduce jobs. Sqoop was retired in 2021.

So I didn’t patch the old text. Posts on ideas that still matter were rewritten for the way teams work now. Posts built around one 2013 tool now teach the idea behind today’s platforms. Hadoop and Hive are still maintained, and Sqoop is retired. Most new platforms are built differently, so these names show up only as a line of history.

What Stayed the Same

The ideas held up better than the tools. Data still outgrows one machine. Work still gets split into pieces, moved by key and added up again. The relational database still sits at the center of most businesses, and it still feeds most analytics platforms.

That’s why MapReduce kept its place in the set, with a new diagram you can follow word by word. It’s also why the database posts are still here. NoSQL, NewSQL, document stores and graph models make more sense once you’ve seen what a relational engine does well.

Many posts now include a small T-SQL demo you can run on SQL Server 2025. Columnstore, the json data type, graph tables and change data capture all live inside one engine you already know. You don’t need a cluster to see most of these ideas work.

How to Read It

The map below groups the 23 posts into five stages. Start at the left and move right, or jump to the stage you need. Each stage builds on the one before it, but every post stands on its own.

Diagram of the Big Data reading map: 23 posts in five stages, the foundations, where the data lives, moving and processing, the databases behind it, and using the data, numbered in reading order from What Is Big Data to the interview answer.

Stage 1: The Foundations

These four posts give you the vocabulary. Read them first if the field is new to you.

Stage 2: Where the Data Lives

Storage changed more than anything else since 2013. These posts explain where analytic data sits today and why.

Stage 3: How Data Moves and Gets Processed

Once you know where data lives, the next question is how it gets there and how the work is split.

Stage 4: The Databases Behind It

This stage covers the operational side. It starts with the relational database and then walks through the other data models.

Stage 5: Using the Data

Storing and moving data has one purpose: better decisions. The last stage is about getting answers you can trust.

Who This Path Is For

The posts assume you know SQL and have worked with a relational database. You don’t need any Hadoop or Spark experience. Every term gets explained the first time it shows up. Almost every post has a diagram that walks through one small example.

If you’re a DBA, stages 2 and 4 will feel closest to home. If you’re a developer moving toward data engineering, start with stage 3. If you manage a data team, read the architecture post and the governance post. Together they give you the whole picture.

What You Can Run Yourself

Twelve posts include a T-SQL demo that runs on SQL Server 2025. The free Developer edition is enough for all of them. Each demo creates its own small database and drops it at the end. Run them on a test server, never on production.

The demos cover columnstore storage, the json data type and graph tables. Others show distances with geography, change data capture, window functions and data quality checks. The other posts teach ideas you can’t run on one server, such as Spark or a distributed SQL database. For those, Big Data Upgraded leans on the diagrams and on small worked examples you can follow on paper.

If You Only Have an Hour

Read four posts. Start with What Is Big Data for the core problem. Then read the architecture post for the whole platform in one picture. MapReduce shows how the work gets split, and the data lake post explains where the results land.

You could argue that nobody needs this list anymore. A managed cloud service hides the storage, the shuffle and the file formats. You click a few buttons, and the queries run. That’s true until a query runs slowly or costs too much. Then the person who knows what happens underneath fixes it in an afternoon.

What to Remember

Big Data Upgraded keeps the ideas and swaps the tools. Most new platforms moved past the 2013 tool list, but the ideas came along. Splitting work, moving data by key and storing it by column all still apply. So do picking the right data model and checking that the numbers are right.

If you read the original series years ago, these posts will feel familiar and new at the same time. If you’re starting today, follow the stages in order. You’ll come out with the vocabulary and with the reasons behind it.

Big data is not a list of tools, it is a set of ideas that outlast any one tool.

Published by Pinal Dave on SQLAuthority. More of my work at pinaldave.com.


Discover more from SQL Authority with Pinal Dave

Subscribe to get the latest posts sent to your email.

Cloud Computing, Data Warehousing, Database, NoSQL
Previous Post
Batch vs Stream Processing: When Data Cannot Wait
Next Post
Measuring Space Used by Tables per Filegroup in T-SQL

Related Posts

6 Comments. Leave new

  • Hi,
    I follows your series of posts regarding the big data.
    but nowhere there is a post regarding the pro and cons, how to implement it, the type of competencies to implement it / use it (DBA, programmer, analyst?);
    also, what are the standards?
    there is a bunch of tools, none of them appear to have a common simple language like SQL or MDX.
    What’s the cost of using this compared to traditional structure data “only”? (so structure/transform sooner (traditional DW) vs later (big data))
    How the big data can fit the self service BI need and trend? (most of the business analysts have problems using structure data, so I cant imagine what’s unstructured data can be for them….)
    how to handle the data quality issues? its fun to gather facebook, twitter and other information, but what’s the purpose if there is no quality in the data? is it the responsibility of each developer consuming these data to do and redo the job? how to compare this to traditional ETL processes? (remember: bad data quality = bad decisions!!!)

    Personally I’m a little lost on where to put the big data in an enterprise. My background is DW oriented, with huge data quality and integrity issues, so structure data is a must have.

    but it was a long and great job!
    continue like this :)

    Reply
  • Personally i feel no data is a bad data. I would call the process of deriving value out of data as a science in itself (Data Science). We have done projects where in data crawling was done from various social websites and output was generated which was meaningful. Regarding hosting of Big Data in an enterprise, we went with setting up a private cloud powered by Openstack and installed hadoop framework.

    Reply
  • Great article, i liked this article as it gives simple definition of very commonly used words like PIG, HIVE, HADOOP. I think big data in combination with IBM Watson is going to be very big thing in future.

    Reply
  • santosh kumar dubey (@dubeysantosh)
    February 5, 2014 6:51 am

    Great article, i liked this article as it gives simple definition of very commonly used words like PIG, HIVE, HADOOP. I think big data in combination with IBM Watson is going to be very big thing in future.

    Reply
  • hi
    thank you 4 your Big Data-serial article
    many useful information!

    Reply
  • I must say a big thank you to you for the post. But i have a question, can you show an example of a working model that transform unstructured data to structured data?. Thank you

    Reply

Leave a Reply

Your email address will not be published. Required fields are marked *

Fill out this field
Fill out this field
Please enter a valid email address.