Introduction
In this skill, we’re focused on what the CompTIA Data+ certification is really built around: the modern data analyst. Analytics today is less about raw tables and more about translating messy data into decisions, especially in a world shaped by cloud platforms, AI, and complex governance requirements. We will use the Data+ domains as a way to understand how analytical work flows from context to insight, and why those steps matter, especially as automation and AI reshape how decisions are made.
Knowledge Check
What is the primary role of a modern data analyst?
The Golden Age of Analytics
In this video, we will step back and look at why analytics has become so central to how organizations make decisions today. The focus is on understanding the broader shift that has elevated the role of data analysts, and why this context matters for how Data+ frames questions and scenarios. This perspective helps set expectations for the kinds of thinking the certification is designed to assess.
Knowledge Check
What is the foundational slab that supports the analytics temple, ensuring that data analytics processes do not become chaotic?
Getting Started in Google Colab
In this video, we will get comfortable with Google Colab as a simple way to follow along with the hands on parts of this skill. We will use a small, realistic dataset to practice the early stages of analyst work, and you will see how the same process can translate whether you prefer notebooks or spreadsheets. The goal is to make the workflow feel familiar and approachable before we go deeper into analysis and communication.
Colab Notebook:
Knowledge Check
When working with two datasets that will be combined later, which action is the best first step to reduce the risk of losing records or creating incorrect matches?
Communicating Data Insights
In this video, we will show results in a way other businesses and people can actually use. Being a data analyst is about deciding how to present that answer so patterns, tradeoffs, and priorities are clear at a glance. The emphasis is on selecting the right format, based on which data points need to be understood and acted on.
Colab Notebook:
Knowledge Check
What is the primary difference between a report and a dashboard?
Core Analytics Techniques and AI Context
In this video, we will look at the different types of analytical techniques analysts use depending on the question being asked, and where those approaches fit within Data+. We will also introduce how modern analytics intersects with AI, clearing up how techniques like machine learning differ from automation, and why governance is the foundation beneath all of it.
Knowledge Check
Which of the following is a subset of AI that uses neural networks, especially with images, audio, and large-scale text?
Challenge 🎉
It's time to check your knowledge before moving on to the next set of videos. Answer the questions below and if you get any of them wrong, review the corresponding video(s) above.
Knowledge Check
In the “temple” mental model, what are the three pillars?
Knowledge Check
Which stage commonly consumes the most time due to missing values, duplicates, inconsistent formats?
Knowledge Check
“Which customers are most likely to churn next month?”
Knowledge Check
Which relationship best matches the AI mental model? (⊃ = subset)
Knowledge Check
A bot logs into systems nightly, pulls data, refreshes a spreadsheet, generates a routine report automatically.
View Transcript
Introduction
0:00Hello and welcome back to the next skill in the CompTIA Data Plus prep course. In the previous
0:06skill, we built the roadmap for the exam, how it's structured, how it's scored, and how to study
0:12smart using science-backed techniques. In skill two, we'll zoom in on the person this certificate
0:19is really built around, the modern data analyst. So what does that really mean in 2026? What about
0:26AI? We've got a lot to unpack in this skill where we explain the modern data analytics roles. We're
0:34living in what many call the golden age of analytics. Organizations are literally swimming
0:40in data, storage is cheap, and competing power is everywhere. And it's also rentable in the cloud,
0:48which means leadership knows there's real money hiding inside of those numbers. But raw tables
0:54don't create decisions. Raw data by itself does not have value. And that's precisely where the
1:01translation layer comes in, turning messy data into usable insight. That's the analyst's job.
1:07So you might be asking, well, why does this matter for the CompTIA Data Plus prep? Here's why. Data
1:14Plus' five domains aren't random exam buckets. And what makes this really easy to approach
1:21educationally is that all those domains are more or less the steps for data analysis. And that's what
1:29we're going to leverage in this skill, because these five domains are basically a job description
1:35broken into parts. So now going back to this raw data that needs to be converted. First, you need
1:42a development environment. And that's the first domain, understand data environments. Step two is
1:48where you acquire and prepare data. So now we have step one, which is a development environment.
1:54Step two, we have the raw data, and now we're going to prepare it. And step three, which is
2:00analyze it. And you could also do some quick EDA here as well. Now you're starting to turn that raw
2:06data into some kind of value. Then number four, you're going to visualize this. This is super
2:12important, because not everybody understands Python, but they sure do understand visualization.
2:19And that's why it's a universal language, if done right. And then finally, we need to lock this down
2:25with governance, which is now an ever growing field, thanks to AI. And now we have something
2:31called AI governance. It's not just data governance. Now it's AI governance. These are intertwined
2:37because in AI governance, you need to understand how that data is being used to train. So data
2:44governance and AI governance overlap. Super interesting. So what does this all mean? By the
2:51end of this skill, you'll be able to explain what a modern data analyst does end to end, why that
2:58role is growing super fast, even with AI, and how data plus domains map directly to real analyst work.
3:07How does that translate into what we're going to actually cover? Let's break it down. In this skill,
3:13we're building a clear picture of modern data analysts from two angles. First, why the role
3:20exists at this scale right now. And this is really about the golden age of analytics,
3:27forcing analytics to the central part of a business. And then B, we'll explore what analysts
3:34actually do day by day. This is the actual workflow, which includes the techniques and
3:40the tools. We'll walk through the temple mental model. This is a visualization of analytics on
3:48top for three pillars, which includes data storage and compute. And here in the roof,
3:55we can call this analytics. And here we have governance. Then we'll walk the analytics process
4:02end to end and connected to reporting and dashboards. And finally, we'll introduce the
4:08techniques such as descriptive, inferential, predictive, and prescriptive. And this is also
4:15where EDA fits into the picture. Finally, we'll connect modern AI flavored work. So that is the AI
4:23machine learning and deep learning, which includes generative AI, NLP, which stands for natural
4:30language processing. And we'll also cover RPA, which is not AI, but it is part of automation
4:38and people get them confused all the time. So we're going to make sure that these are clearly
4:43delineated and that we understand that they're different for this skill. The second skill,
4:48we're mostly focusing on domain 1.0 data concepts and environments, 3.0 data analysis with a little
4:57bit of overlap into the second, fourth, and fifth domains. But generally we're working with domain
5:03one and three. All right. So that's it for this video. And next up, we have the golden age of
5:09analytics and why analytics are in demand. See you there.
The Golden Age of Analytics
0:00Welcome back. So what is the golden age of analytics? The core idea here is simple.
0:06Organizations now treat data as a strategic asset and not just exhaust from operations,
0:14which was the case for quite some time. As a strategic asset, these organizations are also
0:20investing in dashboards and analytics programs and data teams because leaders expect answers like
0:28where should we cut costs or which customers are likely to churn, but most importantly,
0:34they need so with evidence. And that's why modern analysts are in high demand. So companies start
0:41with this raw material or that raw data, but they need people who can translate it into decisions.
0:48So you have that raw data that is now becoming valuable because leaders want to make evidence
0:54based decisions. And to do that, they need people to pull the right fields, clean the messy data,
1:01do the analysis, build visuals, and explain the story all in plain language. For data plus this
1:08context matters because CompTIA questions are usually framed like workplace decisions. That
1:16means given this situation, what would a competent analyst do? And if you understand the role,
1:23the exam stops feeling like random trivia and starts feeling like job scenarios. And now we can
1:29talk about the analytics temple. I guess we could say that this is a Greek style temple.
1:35And you can see that the roof here is going to be analytics and it's being supported by three
1:41pillars. We have data storage and compute or computing power, but no building is complete
1:48without a strong foundation. And that means that the whole thing sits on a foundational slab
1:55labeled governance. So this picture explains precisely why analytics is exploding right now
2:02instead of decades ago. And what's super interesting about this is that right before AI,
2:08that's when everyone is really rushing to use all of the data resources. And that's why data
2:15analytics is also very in demand. Before that it was in demand, but for larger companies. Now
2:21everybody's jumping on the bandwagon because data is the new gold. And since AI is here, well,
2:28what does AI run on? Data. Data has exploded. Transactions, customer interactions, sensor logs,
2:36web events, photos, videos, audio, and storage used to be painfully expensive. Now it's cheap
2:43enough to keep far more historical data for far longer. And that means that analysts can now spot
2:50trends and long tail behavior, but also computing power has skyrocketed and also cloud, which is a
2:58new focus of the CompTIA Data Plus. And that means that analytics aren't just limited to supercomputers.
3:04You can rent serious horsepower on demand through the cloud. And in fact, that's what we're doing
3:10with Google Colab. I have access to tons of GPUs through Google Colab. I do not have supercomputers
3:17here at my house. I pretty much rent all of my compute through the cloud. I do have machines
3:23locally for learning and building local models and things like that, but the heavy horsepower,
3:29I rely on the cloud. And don't forget this foundation of governance. That's what keeps
3:35all that power from becoming chaos because we have access control, privacy, quality standards,
3:43and rules that are responsible for use. Now let's do a real analyst workflow on a churn scenario.
3:51If we decide to zoom into the day-to-day, the data analytics process has five stages. And what I was
3:58super excited about is that, well, these five stages are also the five stages of data analytics
4:05at your job. So that end-to-end analytics process are mirrored in the five domains of Data Plus,
4:12which is convenient and probably by design. All right. So we start with acquisition,
4:18which is domain two, because we don't need to focus on domain one, which is our development
4:23environment. Acquisition means pulling from internal and external sources, but also consider
4:30reality, cleaning and manipulation. Then reality hits because you also have to contend with
4:36cleaning and data manipulation. You have missing values, invalid codes, duplicates, inconsistent
4:43formats, and all of this could consume a huge share of project time. And this is where the
4:50analyst earns their money, really, because you can make data usable and trustworthy. And that's
4:56why cleaning and manipulation takes so much time. And sure, because it offers the most value. And
5:02again, I could just say that this is where the analyst really earns their money because they're
5:07making the data usable and trustworthy. Then we move into analyzing. And this is often where we
5:14start with EDA or exploratory data analysis. You can run simple summaries and visuals to help you
5:21better understand distributions, maybe spot some outliers or form a general hypothesis. Then we
5:29move into visualization and reporting and communication because we have to turn those
5:35results into something humans can act on. And crucially, the process is iterative. That means
5:42that once we're done here, we'll go all the way back if needed to the first step, which means
5:49getting new data and then cleaning it, analyzing it, visualizing it, and then communicating those
5:54results via a report or dashboard, some form of communication. All right, so enough theory.
6:01In the next video, we're going to dive into Google Colab and try this out. We're going to perform an
6:07end-to-end analytics process, much like you would find at a job. See you there.
Getting Started in Google Colab
0:00Welcome back.
0:01In this video, we're going to talk about Google Colab, which is exactly what you see here.
0:06The one that you're going to open looks like this.
0:09It has a bunch of information or code here, but I do want to explain what we're doing
0:13so that this is not confusing.
0:15All this allows us to do is to use Google in the cloud and we have Python loaded here.
0:21And if I go to runtime and then I go to change runtime type, I could choose different GPUs.
0:27I have a CPU here, but I can also use a T4 or V5E-1 or any of these like H100, but you
0:35have to pay for these.
0:36They do let you use these for free, which is pretty cool.
0:40And you can also use R and Julia, not just Python.
0:43So in short, this is probably the easiest way to get Python up and running on your machine.
0:49So I can prove it to you.
0:50I can just say print.
0:51This is what programmers would use when they're testing out a development environment to make
0:56sure that they have Python loaded.
0:58And it takes a second to connect to the cloud, but there it is.
1:01So Python's version.
1:03But to do this, we're going to need the exclamation mark so we can talk to the backend.
1:07And that lets us know that this is Python 3.12.12.
1:11But don't worry.
1:12If you don't know Python, you definitely are going to learn some Python because I'm going
1:17to make it really intuitive.
1:19But more importantly, I'm going to show you that you don't have to use Google Colab.
1:23You can use whatever tool that you're more familiar with.
1:25For example, let's say you work with spreadsheets and that's easier for you.
1:29Let me show you how you can use those.
1:31All right.
1:32So now that you have Google Colab, all you really need is a Google account.
1:35You can see that I'm signed in here in the top right.
1:38And once you're logged in, all you have to do is click on this Google Docs link and it'll
1:43open up and you should be able to see everything that I see and you should be able to run everything
1:49that I run.
1:50However, you will not be able to make changes because this is not yours.
1:54But if you want to make a copy of this, all you do is go to file, save a copy and drive
1:59and then it'll copy this notebook for you.
2:02I accidentally clicked on text, which I'm glad I did.
2:05But if you do that, you can click on delete here.
2:08And if you want to run these code cells, you can do that with this play button or you can
2:12use shift enter.
2:14But if you hover right here in this center area, you'll get either code so that you could
2:18write code or text so you could write text.
2:21I'm going to delete these two and I'm going to show you how to run this first cell.
2:26So shift enter.
2:27And this is talking to the cloud, so it's going to run and then you'll see this check
2:31mark.
2:32It completed in zero seconds.
2:33It was very, very fast.
2:35And before I show you how you can use this in a spreadsheet, we have to load these sources
2:40because in real work, you almost never start with a perfect table.
2:44I wish you start with files from different systems.
2:48So here we're going to use our CRM and our demo file.
2:53So this CRM right here is an internal CRM that has churn and some interesting behavior.
3:00And we can call this the internal truth.
3:02And then here we have a kind of a demo.
3:05I call this demo and this is just enrichment.
3:08This is external data, but there is data mismatch and duplicates and missing data because I
3:13want it to be a little bit realistic.
3:16So let's click on this folder here and you're going to see that if I click on refresh, you
3:21just have sample data and all of this.
3:23Not really a big deal.
3:24But once I run this, we're loading the data.
3:27So this is the very first part.
3:29So now that you've loaded this data and you click on refresh, now you have CRM customers
3:34CSV, which is a comma separated value.
3:38And now you have demographics dot CSV.
3:40That's the demo, which is the enrichment data, which is dot CSV as well.
3:45So you can just click on these and then download these.
3:47So once you've downloaded these, you can go to Excel or even Google Sheets, go to file
3:53and then import.
3:55And once you do that, then you would go to upload here and then browse or drag a file.
4:00And then once you uploaded that, now you would have that same CSV for the CRM, which is,
4:06I closed this one.
4:07We don't need that, which is here.
4:08OK, so this is the CRM data.
4:10So you don't have to use Google Colab.
4:13You can use spreadsheets.
4:14So if this is more familiar for you, you can see that we have customer ID from 1001 to
4:201009.
4:21If we go back to this data here, you can see that we have 1001 to 1010.
4:26And we're going to explain that in just a second.
4:28Why do we have 1010, but there's no 1010 here?
4:32This is going to come up and this is part of the lesson.
4:35OK.
4:36So now that you have these two data sets loaded into a table, you can convert to a table like
4:41this.
4:42And I'm going to close this.
4:43And now you can do things in this table.
4:45So I can do math if you have the correct data types.
4:48And again, that's part of cleaning.
4:50So let's go back over here and get started.
4:53All right.
4:54So all I've done is just imported pandas, which is a digital version or let's say a
4:58Python version.
4:59Right.
5:00It's a Python package for tables and data manipulation.
5:04And then we have NumPy, which because pandas is built on top of NumPy.
5:09And then you have matplotlib for visualizations.
5:12Normally, Colab has these installed, but I'd like to show this as a best practice in case
5:17you're using non-Colab like Jupyter Notebook, which is a local version of what you see here
5:23or a non-cloud version.
5:24All right.
5:25So don't worry about all of this stuff.
5:27We're just loading the data.
5:29That's all we're really doing.
5:30And I'm giving you some options so that you can use spreadsheets or any other tool by
5:36downloading the CSV and performing those actions.
5:40Because this is a vendor neutral certification, which means, well, if you understand the process
5:46on multiple tools, that's all you really need to do.
5:49So you can follow along with me and it won't be too difficult to do that.
5:52We're not writing any code.
5:54You can use Google Colab or you can use the spreadsheet.
5:57I'll leave it up to you.
5:58All right.
5:59So you've loaded this data or you have it in a spreadsheet.
6:02So now I'm going to return the first five rows with this code right here.
6:07And now you can see that we have 1001 to 1005.
6:11But let's say you want to look at the last five, you can do something called tail.
6:15So the last five here.
6:17And you can see that we have 1010 to 1004.
6:21Now let's do the same with the CRM tail.
6:24OK.
6:25And you see this one goes to customer ID 1009.
6:28It doesn't go to 1010.
6:30So that is just a little foreshadowing of something that we'll need to address together.
6:34All right.
6:35This is pretty much what data acquisition actually feels like.
6:39So where are we in this process?
6:41We basically are in the first stage, which is acquisition.
6:45And this is before you clean anything.
6:47You can confirm what you actually have.
6:49And that's what we've done here with head and tail.
6:51I could also just do CRM like this and get the first 10.
6:56And in this case, it's trying to show me as much as it can.
6:59It will truncate it if this is a big data set.
7:02And so the first thing that I would do is to inspect the columns and the row counts.
7:07So the columns is going to tell you what fields exist.
7:10So here are the columns.
7:11So you have customer ID, plan, tier, churn, monthly, and so on.
7:17And if you see some missing values, then you'll know what kind of values are missing.
7:21Row count tells you if you're about to lose data during a join.
7:26So sometimes you can perform a join where you take two tables and combine them.
7:31And based on that join, sometimes you can lose rows.
7:35And looking at the row count is a good way to make sure that you're staying one step
7:40ahead of that, meaning you're not losing any rows during a join.
7:44All right.
7:45So this is the habit.
7:47You want to always check the columns, number of rows, and use a tail or head or just look
7:54at the data like this.
7:56And you can also do so here in a spreadsheet.
7:58All right.
7:59Now let's move down over here.
8:00So we've already acquired and loaded that.
8:03So let's, well, we did that here inside of Colab, but we need to load those here.
8:08And the way that you would do that, let's say that we didn't use pandas, but instead
8:13we have them here, right?
8:15And if that's the case, you can upload CSVs to Google Colab.
8:19And if you run this, then this is the load and inspect.
8:23And then here, it's going to tell me what the column names are and the row counts for
8:29each data set.
8:30So I have 10 and then eight, oh, oh, warning.
8:32Let's make sure that we keep this in mind.
8:35And whenever we perform joins, we want to make sure that we're not losing rows.
8:38And now here's where we make our money.
8:40This is the stage where analysts actually earn their paycheck and we're making the data
8:46set joinable and trustworthy.
8:49So what does that mean?
8:50So first, let's go ahead and deal with this one right here.
8:53So when you're cleaning and manipulating, that includes fixing things like white space,
8:58which are hard to see, and also data type mismatch.
9:02So before we do this, let me just show you what this is here.
9:05Let's inspect this a little bit more.
9:07Let's say crm.info.
9:09OK, so what we want to do is make sure that the CRM customer ID and the demo or demographics
9:15customer ID are both going to be strings, because right now it's an integer.
9:20So we want to change that to string.
9:22And we also want to fix the white space, remove the white space, because when you're doing
9:26a join, it's very strict.
9:28If you have one extra space, there's going to be no match, meaning if you had one extra
9:33space, we would get a bunch of unknown regions and the join would fail.
9:38So that's why we're doing this.
9:40So let's go back to here.
9:41Now we're going to see that they're integers and we're going to run.
9:45Let's go ahead and just run this here so that you can see the changes.
9:49So now I'm going to run this cell again so that I can see what's changed.
9:54And now I'm going to be paying attention to customer underscore ID here.
9:57And now you can see object and object is a string.
10:01That's just what Panda calls a string.
10:03So just have to memorize that if you're working with Pandas.
10:07And if you're working with a spreadsheet, it's going to be slightly different.
10:10But what we're doing is just changing it from an integer or a number into text or a string,
10:16which Pandas calls an object.
10:17And if I change this to demo, we would see that we also have now the customer ID as an
10:22object.
10:23OK, so I'm just going to delete this because we're going to do this down here again.
10:28All right.
10:29So next you want to drop duplicates.
10:31What you're doing is you're collapsing multiple records to one per customer ID.
10:36And here's something important to consider.
10:39In real work, you don't blindly keep the first row right here.
10:43And I put this here because we're just doing that because this is a toy demo.
10:47You would decide that based on the data that you're working with, maybe the latest update
10:52or the most complete row, et cetera.
10:54In this case, it's a toy data set.
10:56So we're just going to keep the first row like this.
10:59So that's something to consider.
11:00And then here we're going to merge before we do merge the duplicates.
11:04The reason that we're doing this, too, is that duplicates can inflate counts and distort
11:09rates.
11:10So this is important.
11:11Now we're going to do a merge, which is a join.
11:13So here we're going to emphasize the left join.
11:17I put that here in the comments, but you can see here how is a left join.
11:21So our CRM is the analysis population and demo is the optional enrichment, the demographic.
11:29And one thing to notice here is that we're not using demo as the base.
11:33And as a result, if we go back up here, you can see that we have 10, 10 here and demo.
11:39And since we're not using 10, 10 or demo as the base, that means that it's going to disappear
11:43here when we perform that join.
11:45Let's go over here.
11:46OK, so that explains that join there.
11:48And then finally, we have simple missing handling.
11:52And for that, we're just going to handle missing values with fill and a unknown.
11:57And that's how we do that right here.
11:59We're just going to handle the missing values.
12:01And when we see a known, we're not saying that the region is truly unknown.
12:06What we're saying is that we're labeling missing data so that we can track it.
12:11That's the whole idea here.
12:12All right.
12:13So now I'm going to grab this little line here and remove it.
12:16And then I'm going to run this cell and paste this here.
12:19And this is where you check everything, right?
12:20We can do this one by one as well.
12:22So here we have the rows after dupe.
12:25Let's do this one here.
12:26OK, so nine.
12:27That's good.
12:28And then finally, we're going to check for the percentage of missing demographics.
12:34And here we have 33.3 missing.
12:36And then finally, we can do this, which is merged head.
12:40And now you can see that you have both of those tables merged.
12:43So what's important to consider here is that every analyst should run sanity checks after
12:48a merge.
12:49And that's what these were right here.
12:50So you want to check your work, right?
12:53If you do a join in these spreadsheets, you want to do the same thing because this is
12:58vendor neutral.
12:59We just need to do the sanity checks after you do the join.
13:03And now we can finally analyze.
13:06But we start simple.
13:07We're going to use descriptive statistics first and a little bit of EDA, but not too
13:13much.
13:14So we want to analyze the churn rate by plant tier.
13:17And we're going to say churn by tier.
13:20And that's how we get churn rate by plant tier.
13:24And then here's a very quick EDA.
13:25And what we're doing here is just asking questions.
13:28Do the ranges look same?
13:30Do we see weird minimums or maximums, right?
13:34Or is the data skewed in some particular way?
13:37So again, here for the EDA, we're just asking questions, not making predictions about churn
13:42just yet.
13:43You're just identifying patterns that are worth investigating.
13:46So let's go ahead and run this.
13:48So we have basic and enterprise at 50% churn, and then we have pro at 33.
13:55Here's some very simple statistics.
13:57I think this is dot describe.
13:59Yep.
14:00This is how we do that.
14:01And here we can create a simple visualization, which is, in essence, the same visualization
14:07or the same insight, I should say, as the table, but faster for humans to scan because
14:13it's visualization and not all this code.
14:16So let's go ahead and run this.
14:17All right.
14:18Let's check this out.
14:19So here it's the same information, right?
14:2050%, 50%, 33%.
14:23But the cool part here with this, bar charts are great for category comparisons.
14:28Look at that.
14:29You can see that and you can show this to non-technical audience and you can add labels
14:34and titles, whatever you need for interoperability.
14:37Okay.
14:38So here, let's do a quick distribution check on last login days.
14:43So what we're doing here is pretty classic EDA.
14:46How is the engagement distributed?
14:49All right.
14:50So histograms will show you a shape like, and that means like a cluster or let me just
14:56show you.
14:57All right.
14:58So if you had a normal distribution, you would see something that goes up like this, right?
15:02Let me change the pen color.
15:04A normal distribution.
15:05So again, here's a different color.
15:07A normal distribution might be clumped up like this in the middle, but you can see that
15:13this distribution is skewed like this and this is called a right skew.
15:17So now we're looking at the distribution and again, we're just looking at the shape and
15:21this supports hypothesis building.
15:23You can say our most inactive customers churning and other questions.
15:28And you can think of this as a investigation, like a quick investigations.
15:33That's sort of like what I like to think of EDA.
15:35And if we go down here, finally we have some kind of communication, right?
15:41So we're not going to get too deep into this, but if we run this, here's actionable takeaway.
15:45The highest churn is in the basic tier.
15:47Well, it is, but it is also in the highest, in the, where is it?
15:52Let's go here.
15:53It's also an enterprise tier.
15:55They're both 50% and next step segment that tier by inactivity and then some kind of next
16:01steps.
16:02All right.
16:03So this is pretty cool.
16:04We acquired and loaded that CRM CSV, it was multiple steps, right?
16:08And then we clean and manipulate.
16:10We fix IDs dedup, meaning deduplicated.
16:14We merged and we label missingness as unknown.
16:18Here we're going to clean and manipulate like this, what we just talked about.
16:22And then finally analyze.
16:23We compute churn rate and run some kind of a quick EDA.
16:27And then over here, we're going to check out the distribution.
16:31And this is also just the histogram.
16:33That's what it's called.
16:34And then some kind of communication at the end.
16:37This is just to show you how this works.
16:39And since I didn't go into it very detailed, let's do that in the next video where we talk
16:44about communication and insights.
16:47And to do that, we'll look at tables and visuals and reports and dashboards.
16:52See you there.
Communicating Data Insights
0:00A big part of being a modern data analyst is communication or the skill of communicating
0:06insights.
0:07You could say that, okay, we're choosing tables, visuals, reports, and dashboards and
0:12call it a day.
0:13But when we're looking at tables, we know that they are precise, but they're hard for
0:17humans to scan for patterns.
0:20They're precise, but there is a higher cognitive effort, meaning that a chart often makes the
0:25same point almost effortlessly and instantly.
0:29You can spot trends, outliers, and comparisons.
0:34Then we package visuals into deliverables.
0:37A report is a point in time snapshot because these charts and maps not only give those
0:43patterns visible at a glance, but a lower cognitive load.
0:47So a report is visuals plus narrative that explains what happened and why it matters.
0:53A dashboard, on the other hand, is for ongoing monitoring, meaning visuals that are updated
0:59regularly and support filtering and drill down.
1:03And be prepared for a scenario where you're given an audience and a goal and which deliverables
1:10and visualization choices might be best.
1:13But we really want to look at these deliverables and examine reports and dashboards with a
1:18little more detail.
1:20So now let's go back to Colab and try this out.
1:23Just remember, this is about communication.
1:26Analysts don't just compute answers, they package answers.
1:30That's a good way to think of it.
1:32Now in this Colab notebook, we're going to take the same churn insight and deliver it
1:37as a table, a chart, a short report, and a dashboard like view.
1:43Again, we're not going to be able to go into super detail here because we're just doing
1:48an overview of how all of these work together in the role of a modern data analyst.
1:53All right.
1:54So we've run our imports and now we need to load or recreate the merged data set because
2:00we already merged it.
2:01And so this should be pretty familiar.
2:03So I think that we can go over here to this folder.
2:06You'll see that there are no files here.
2:08So if you need those files again, you're going to need to go over here and download them
2:12from this place because I made them downloadable here.
2:16I turned them into CSVs just for this exercise.
2:19So definitely go to demo one.
2:21You should still have these downloaded if that's what you're doing.
2:23I'll close this and now we can see our table and the way that we're doing this is DF head.
2:30So let's go through this together.
2:31And a quick note, this is domain four style of thinking.
2:35You're choosing the best deliverable for the audience and goal.
2:38And again, we have two data sets, two different worlds.
2:41We have the CRM, which is the internal system.
2:44And then we have the demographics, which is this third party external enrichment.
2:50This could be region, age, income, and so on.
2:53And all of this mirrors a real business workflow.
2:57You want to join business behavior with some kind of context.
3:01And as you see here, we're also taking care of the white space because otherwise if the
3:06customer ID is numeric, like an integer, the join could fail silently because the keys
3:12don't match.
3:13And that's another important reason why we're doing that.
3:17Here we're dropping duplicates and here we're doing that join.
3:20And we're using fill NA for those missing records.
3:23And we're saying that we don't know what those missing values are.
3:26We're not saying anything specifically like unknown age.
3:30We don't know what age it is.
3:31It's just more like cracking the missing value, right?
3:34Okay.
3:35So now we're ready for that merge.
3:36Okay.
3:37So here's our table.
3:38So we have a pivot table.
3:39And what we want to do here is keep all the CRM customers and pull in the demographics
3:44information when available.
3:46And that's why we're also going to do, let's go back up here.
3:49I believe the join is here.
3:50There's the join, a left join.
3:52This means that the left is the CRM, which is the truth set, meaning the customers we
3:58have churn labels for.
4:01Demographics is the optional data, the enrichment.
4:05Okay.
4:06So let's go ahead and run this first table here.
4:08We could also say this is the table version, meaning that we have precise compact and we're
4:14aggregating raw rows into decision ready metrics.
4:18All right.
4:19So looking at this first pivot, this looks pretty familiar.
4:22And here we're also sorting the values and this makes the story, I guess, more readable.
4:28You're sorting and renaming.
4:30And this is something you could consider this reporting polish, right?
4:33To make it look good.
4:34And what we're trying to convey here is which tier churns the most.
4:38So let's do the second one here.
4:40And here we want to shape it like a dashboard matrix.
4:43Okay.
4:44This is more of a report place in time.
4:46So let's go ahead and do this one here.
4:48It's categories on the side, and then you have these segments across the top.
4:52And so here the index is now the plan tier and the columns are the different regions.
4:57Okay.
4:58And this creates a grid that's going to be visually easier to support comparisons.
5:02And that's why we're doing this.
5:04And you could also say this is where the raw data becomes a monitoring view.
5:09And if you're working in a spreadsheet, well, pivot tables are the backbone of Excel and
5:13dashboards and BI summary.
5:16So this should be very easy and other tools as well.
5:18Okay.
5:19Now going down to this pivot chart here.
5:21And like we said before, these tables are exact, but charts are fast for human brains.
5:27So let's go ahead and run this.
5:28All right.
5:29So it's really easy to see that we have the highest return rates are both basic and enterprise,
5:34nothing new.
5:35And now I want to compare a report versus a dashboard.
5:38This is really important.
5:40And I'm going to leave these here for you.
5:41A report is a point in time and a dashboard is ongoing.
5:46That's the takeaway.
5:47So when we're talking about a report, we really mean narrative bullets and some kind of chart
5:53or table or a snapshot.
5:55So you could say that churn is concentrated around the less engaged customers, but you
6:02actually haven't proven it in this notebook.
6:04So it really depends.
6:06This is toy data, so we can't take it too far.
6:10That would be more of a hypothesis that you'd need to validate with some kind of a quick
6:14check, like a correlation or grouping by last login days or something like that.
6:19So what we're doing here is labeling what's confirmed or comparing that to what's suspected,
6:24right?
6:25So here we're labeling what's confirmed versus what we suspect.
6:30Now let's take a look at the dashboard.
6:31Well, let me run this cell here.
6:33OK, so here's a dashboard.
6:35It's saying here that the highest churn rate is basic, but this is just a demo, right?
6:40You can see that they're tied.
6:41So we get the highest churn and churn appears concentrated among less engaged customers,
6:46right?
6:47And the next step might be target retention offers for high risk customers in the top
6:51churn tier.
6:52OK, cool.
6:53Now let's look at this dashboard.
6:55A dashboard is not a story.
6:57That's the first thing to say.
6:59It's more of an interface, so it's meant for repeated use.
7:03That's the difference between a report and a dashboard.
7:05And that's why we're using a reusable function.
7:09It's more of an interactive approach.
7:11You could say an interactive mindset.
7:14And here we're going to plot churn by tier for a particular region.
7:19And you could say this is simulating kind of like a filter.
7:22It's the same metric, but we're packaging it a little bit different.
7:26It's the same as a pivot table idea, but it's just applied after filtering.
7:31That's one way to think of it.
7:32We're handling empty filters gracefully.
7:35So that's just a little detail.
7:36And then we're going to run two regions live.
7:39So let's go ahead and run this and let me show you what I mean.
7:42And now let's run this.
7:43OK, so now you have basic and pro and then you have basic and enterprise as a comparison.
7:49You can also use this UI that I have right here.
7:52Let me run this and show you how that works.
7:55And you can just click on South.
7:57And this is really great for non-technical audiences.
8:00Here's the unknown, West and Midwest.
8:04So we'll go back to South here.
8:05I'm showing you that you could use this user interface
8:08dropdown for the dashboard.
8:10All right, so that's it for this video.
8:12And I'll see you in the next video on core analytics techniques
8:16and where EDA fits in.
8:18See you there.
Core Analytics Techniques and AI Context
0:00Under the hood, analytics uses different classes of techniques,
0:04depending on the question that could be descriptive, inferential,
0:08predictive, and prescriptive.
0:10And that's what we're going to cover in this video.
0:12We're going to connect directly to the data plus domain 3.0 analysis,
0:18especially choosing the right method for the situation. And we'll start.
0:21So let's start with descriptive here.
0:24We're answering what happened and we can use summaries and distributions,
0:29but this is where the EDA usually starts.
0:33When we're looking at inferential statistics,
0:35we're talking about confidence intervals,
0:37which we'll get into a little bit later, including tests and regression.
0:41But for now, inferential answers,
0:43what can we conclude about a population from a sample? Okay.
0:48And when we're talking about predictive analytics,
0:50we're talking about forecasting risk and likelihood.
0:54And then finally here with prescriptive and you can think of a prescription
0:59here. We're recommending actions under constraints,
1:02take two cookies and call me in the morning. That's my kind of doctor,
1:07but that's just really an easy way to remember this. And again,
1:10we'll have a lot more opportunities on how to choose the right method for the
1:14question.
1:15When we continue to explore core analytics techniques throughout this course,
1:20and we're just getting started. All right. So now to talk about AI here,
1:24modern analytics, isn't just charts and averages anymore.
1:28And it's increasingly AI powered. And when I say AI, AI,
1:32isn't just one thing. So here AI is an umbrella term for machine learning,
1:36deep learning, generative AI, but it's not RPA.
1:42Then that's what I wanted to talk about. Artificial intelligence.
1:45It's really an umbrella term that includes everything such as
1:50deep learning, machine learning, generative AI, NLP, and more.
1:54The purpose of this slide is to let you know that RPA is not
1:59the same as AI. RPA stands for robotic process automation.
2:04Let's break this down. So when you're thinking about AI,
2:07just think that AI is an umbrella term.
2:09And there's two main fields, machine learning and deep learning.
2:13These are subsets of AI. Don't worry about that for now.
2:16These are techniques that imitate aspects of human behavior or
2:21decision-making. A simpler way to think of this is software modeled
2:26after the brain.
2:27It uses artificial neurons and all of this really super interesting stuff that
2:32we're going to get into in this course. So by the end of this course,
2:35you'll have a good idea of machine learning and all the subsets and related
2:39technologies. For instance, when we're talking about machine learning,
2:43this is really a subset of AI where algorithms learn patterns from data.
2:48Here, when we're looking at deep learning,
2:50this is a subset of machine learning,
2:52which is a subset of AI that uses neural networks,
2:57especially with images, audio,
3:00and large scale text. And that's the software modeled after the brain.
3:04And that's why we're using neural networks or more accurately,
3:08we should say artificial neural networks.
3:11We have biological neural networks and these models,
3:15AI models use artificial neural networks, and there's so much more.
3:20For example, generative AI,
3:22this is where the chat bots like chat GPT and Grok and Anthropic
3:27come from, which includes NLP. We have here,
3:30which is natural language processing. Again,
3:33we'll get into all of this very soon. Just remember in analyst work,
3:37this can help draft narratives, summarize documents,
3:40suggest SQL code or scaffolding,
3:43or translate findings into plain English,
3:47but you're still going to verify outputs and follow governance.
3:50You don't just rely on AI. It's actually pretty dangerous to do that.
3:54And then finally we have the whole purpose of this slide,
3:57which is to compare AI and RPA, which is again,
4:01robotic process automation. This is where people get confused.
4:06RPA is not learning it's rule-based automation.
4:10These are bots that do repetitive steps like exporting a report,
4:15moving files, refreshing templates, or maybe sending routine outputs.
4:20So just remember, these two are not the same.
4:23RPA is kind of the boring one.
4:26And all these other technologies are the exciting AI ones.
4:29So just remember RPA is kind of automated, a little bit boring,
4:33but that's a good thing because we want it to be automated and rule-based and
4:38predictable.
4:39And when we look at what machine learning looks like every day for an analyst,
4:43you might see something like customer segmentation here.
4:46We're grouping customers to support retention or maybe onboarding offers or
4:51to identify the strongest customers. Anomaly detection here.
4:55You want to look for those spikes, right?
4:57You want to spot unusual spikes and logs errors or activities.
5:01Outliers are a good way to spot those. And finally, we have sales forecasting.
5:06Here we're estimating future demand from historical data or historical
5:11patterns.
5:11And all of this power only works if people trust the results and the
5:16organization stays safe.
5:19That's why governance is the foundation in the temple that we looked at.
5:23Remember that analytics roof, the three pillars.
5:26And then now here we're talking about governance, the foundation.
5:31These include consistent definitions, access control, privacy rules,
5:35quality expectations, and even documentation practices.
5:40So if you don't write documentation, maybe you should,
5:42it's a big part of data governance and AI governance moving forward.
5:47This is what lets analytics scale responsibly. All right.
5:50So here we have 1.0. You want to understand,
5:54this is the first domain. These are systems, data concepts, and AI terms.
5:59All right. So let's map these to the different domains. So for 1.0,
6:04we're really thinking about systems, terms, AI concepts,
6:09data concepts. For 2.0, we're thinking about acquisition and cleaning.
6:14That means acquiring, cleaning, shaping the data, basically preparing it.
6:18And for number three, where we analyze, these are analysis methods,
6:23statistics, machine learning choices, maybe deep learning.
6:26And then once you have something to talk about, well,
6:28maybe you can create some visualization,
6:30some charts or a report or dashboard, depending on the context.
6:35And then finally governance. Here,
6:38we want to ensure access privacy quality and trust for responsible use.
6:43And that's it for this last video,
6:45where we explain the modern data analytics roles. And until next time,
6:50I hope this has been informative. And I'd like to thank you for viewing.
Team training path
Turn this skill into assignable team training
This free skill is a preview of the courses your team can assign, track, and report on with CBT Nuggets.
$708
seat / year