Query Data with DynamoDB
Let's supercharge our DynamoDB skills with queries!
Introduction
⚡️ 30 second Summary
In the last project, you joined the NextWork team as our Data Engineer - so good to have you on board!
We started building a DynamoDB database to store our projects, videos and community activity.
Now that our data is all in one place, can you please investigate what insights we can grab from this database?
In this project, get ready to...
- 🌟 Upload tables and data into DynamoDB using CloudShell
- 🔎 Query DyanmoDB through the console and CloudShell
- 🪄 Set up a transaction i.e. edit TWO tables at once
Let's roll up our sleeves and get these done in the next hour. 💪
Want a complete demo of how to do this project, from start to finish? Check out our 🎬 walkthrough with Natasha 🎬
If you're up for a bit of a challenge, quiz yourself on the key concepts up ahead in this project.
Before we start Step #1...
Before we get started, it's important that you know what we're trying to do today.
Login with your IAM user
For this project you'll need your IAM user, not your root user.
First things first... do you have an IAM user?
No
Oooo it's the start of a new era!
If you don't have an IAM user yet - here are the steps to create one (this takes less than 10 mins).
What is an IAM user? Why are we setting one up?
In AWS, a user is a person or a computer that can do things on the AWS cloud.
When you create an AWS account for the first time, the login you get is called the root user of the AWS account. AWS actually recommends to not use your root user for everyday tasks to protect it from security breaches.
You should create IAM users instead. If a root user is a master key to your AWS account, think of IAM users as key copies. IAM users have separate usernames and passwords to your root user, and you can set them to have limited access to your account's resources.
- Head to your AWS Account as the root user.
- Open the AWS IAM console.
- From the left hand navigation panel, choose Users.
- Choose Create user.
- For the User name, name it:
[[YOURNAME="enter your name"]]-IAM-Admin
- Make sure to select the checkbox next to Provide user access to the AWS Management Console - optional.
Note
This does not apply to all accounts, but if you're prompted with a pop up panel that says Are you providing access to a person?, choose I want to create an IAM user.
- For the console password, choose Custom password.
- Type in a password that you will be able to remember/access in the future.
Top tip
You will use this password for all future projects, so make sure to choose a secure one!
- Deselect the checkbox for Users must create a new password at next sign-in - Recommended.
- Choose Next.
- In the permissions set up page, choose Attach policies directly.
- From the list of Permissions policies, select AdministratorAccess.
- Choose Next.
- Choose Create user.
- Voilà - you've just created your new user! Stay on this page.
- Choose Download .csv file.
- Copy the Console sign-in URL.
- Now you're ready to start using your IAM user. 🏁
- Log out of your root user's AWS Account.
- Paste and go to your copied console sign-in URL.
- Open your downloaded .csv file containing your user's access instructions.
- Log in using your IAM user's username and password in the .csv file.
- Once you're logged in, you're ready to use your IAM user for this project! Make sure to keep the login details safe - you'll need them for the entire 6 Day DevOps Challenge!
Yes
Nice! Log in to the AWS Management Console with your IAM Admin User.
Note
PLEASE make sure you log in to your IAM Admin User instead of the root user - it's truly best practice for account security.
Set Up DynamoDB Tables with AWS CLI
Let's start with our most fundamental ingredient... a DynamoDB table! Once we have this, we can learn how to populate it with data and query it for insights. Lovely!
In this step, get ready to:
- Run commands inside AWS CloudShell to create DynamoDB tables.
Note
Note: If you've never used CloudShell or DynamoDB before, we recommend checking out the previous part of this series to learn more about DynamoDB tables.
- At the top of your AWS Management Console, select the icon for AWS CloudShell.
What is CloudShell?
AWS CloudShell is shell in your AWS Management Console, which means it's a space for you to run code! The awesome thing about AWS CloudShell is that it already has AWS CLI pre-installed.
💡 What is CLI? AWS CLI (Command Line Interface) is a software that lets you create, delete and update AWS resources with commands instead of clicking through your console.
You usually have to install AWS CLI into your computer to use it, but in our case, CloudShell already has CLI installed for us (thank you CloudShell 🙏).
Extra for Experts: As you advance in your cloud engineering career, you'll find that the AWS CLI often becomes your go-to tool. Engineers use the CLI to automate tasks and manage AWS resources efficiently using scripts, making it essential for managing your cloud environment in an efficient way.
While the AWS Management Console is fantastic for learning and having a visual guide, the CLI provides the speed and versatility that professionals need for complex tasks.
- Wait 30 seconds for your environment to be ready.
- Run these commands to create new tables:
aws dynamodb create-table \
--table-name ContentCatalog \
--attribute-definitions \
AttributeName=Id,AttributeType=N \
--key-schema \
AttributeName=Id,KeyType=HASH \
--provisioned-throughput \
ReadCapacityUnits=1,WriteCapacityUnits=1 \
--query "TableDescription.TableStatus"
aws dynamodb create-table \
--table-name Forum \
--attribute-definitions \
AttributeName=Name,AttributeType=S \
--key-schema \
AttributeName=Name,KeyType=HASH \
--provisioned-throughput \
ReadCapacityUnits=1,WriteCapacityUnits=1 \
--query "TableDescription.TableStatus"
aws dynamodb create-table \
--table-name Post \
--attribute-definitions \
AttributeName=ForumName,AttributeType=S \
AttributeName=Subject,AttributeType=S \
--key-schema \
AttributeName=ForumName,KeyType=HASH \
AttributeName=Subject,KeyType=RANGE \
--provisioned-throughput \
ReadCapacityUnits=1,WriteCapacityUnits=1 \
--query "TableDescription.TableStatus"
aws dynamodb create-table \
--table-name Comment \
--attribute-definitions \
AttributeName=Id,AttributeType=S \
AttributeName=CommentDateTime,AttributeType=S \
--key-schema \
AttributeName=Id,KeyType=HASH \
AttributeName=CommentDateTime,KeyType=RANGE \
--provisioned-throughput \
ReadCapacityUnits=1,WriteCapacityUnits=1 \
--query "TableDescription.TableStatus"
What does this command do?
This script includes commands to create four new tables in AWS DynamoDB, each with specific attributes and settings.
The four tables created are:
- ContentCatalog Table: This table has a numeric attribute called Id.
- Forum Table: This table has a partition key called Name.
- Post Table: This table has a partition key called ForumName and a sort key called Subject. We'll dive into sort keys soon!
- Comment Table: This table has a partition key called Id and a sort key called CommentDateTime.
- ✈️ Off we goooooo!! AWS CloudShell tells AWS CLI to work with DynamoDB and create those tables for you.
- Tip: if your terminal stops updating, you might need to press Enter on your keyboard to run the last command.
If a Safe Paste panel appears, select Paste.
Bonus tip: You're going to see a lot of multi-line code in this project!
Untick the checkbox Ask before pasting multi-line code so this panel doesn't pop up every time.
- Let's confirm those tables were actually created. Run these wait commands, and wait until they all run and end:
aws dynamodb wait table-exists --table-name ContentCatalog
aws dynamodb wait table-exists --table-name Forum
aws dynamodb wait table-exists --table-name Post
aws dynamodb wait table-exists --table-name Comment
What are wait commands?
When you run a wait command, you're telling your terminal to keep waiting until a condition is finally met. In our case, we're saying "keep waiting - don't finish running this command until this table has been created."
Wait commands are helpful for making sure necessary resources have been created before you move on. Otherwise, future commands that depend on your resources would automatically fail!
Extra for Experts: Can you guess what might happen if you ran a wait command before creating the resource itself? The wait command will keep running and waiting... eventually it'll decide that the resource doesn't exist and fail i.e. stop!
- Head into your DynamoDB console and select the Tables tab.
What is DynamoDB?
Amazon DynamoDB is a non-relational database service.
Non-relational databases use structures other than rows and columns to organise data.
Extra for Experts: DynamoDB can also be described as a NoSQL database, or a key-value database. A NoSQL database means you would not use SQL to query it, while key-value is a specific way to store data that's flexible and efficient.
💡 What is a DynamoDB table? In Amazon DynamoDB, all data is organized into tables! Unlike relational databases which use rows and columns, DynamoDB tables use items and attributes. Let's see what these items and attributes look like in the next few steps.
💡 So a DynamoDB table is a table... without rows and columns? How is that possible? It might sound impossible to have a table that doesn't use rows and columns, but it's true that DynamoDB tables aren't like the traditional structure!
You'll see a DynamoDB table in action soon, but in short, imagine if you had a table where each row had a different number of fields, and every single cell can have a different column header. Instead of the typical relational database structure (where each row has the same columns and column headers), database tables are a lot more flexible.
- Confirm that you see four new tables!
- Tip: You might need to refresh your page if you don't see them straight away.
Load Data into Your Tables
Now that we've got our DynamoDB tables set up, we can actually start to load data in.
In this step, get ready to:
- Load some data into DynamoDB tables.
- View and update your loaded data.
- Head back into your CloudShell terminal.
- Download and unzip this zip file containing data:
curl -O https://storage.googleapis.com/nextwork_course_resources/courses/aws/AWS%20Project%20People%20projects/Project%3A%20Query%20Data%20with%20DynamoDB/nextworksampledata.zip
unzip nextworksampledata.zip
cd nextworksampledata
We're unzipping a file... into CloudShell?
That's right! CloudShell is an environment that can also handle 1GB of storage, so you can save files inside CloudShell.
And here's a fun tip: you could write and store scripts right in CloudShell to automate repetitive tasks, so you won't need to run AWS CLI commands line by line!
Extra for Experts: CloudShell's storage is persistent, meaning your files and data will stay available across sessions as long as you stay within the 1 GB limit.
🚨 Run into this error?
That's because you've done the previous part of this project, and we haven't deleted these downloaded files from CloudShell last time!
Now, CloudShell wants to know which copy of the same file you'd like to keep.
You can press y on all prompts to use the most recent version of the file.
- run ls to confirm that all files are now inside your CloudShell environment.
- Want to see what's inside these files?
- Run cat Forum.json
What does this command do?
The cat command opens up and lets you read files directly in the terminal.
So when you run cat Forum.json, the terminal will show you all the data inside the Forum.json file right on your screen. This is such a quick way to view or verify the contents of a file without having to open it somewhere else.
Extra for Experts: cat actually stands for concatenate, which is a command for joining files together. But, if you only run cat with a single file, you end up just reading it instead!
So what's inside Forum.json?
Let's break it down! Forum.json contains data that's been formatted specifically for loading into DynamoDB:
- "Forum": tells DynamoDB that the data relates to the Forum table.
- "PutRequest": tells DynamoDB to add a new item into Forum table.
- Then the rest of each PutRequest includes the attributes of this new item! Notice how the first item has five attributes (Name, Category, Posts, Comments, Views) but the second item only has three.
- This is a great example of a DynamoDB table's flexibility - every item can have any number of attributes and attribute titles.
- If we were using a relational database... then both items would need to have a value for all the attribute names in the table! This makes your database bigger (and slower).
- Load the data of all four files into DynamoDB using AWS CLI's batch-write-item command:
aws dynamodb batch-write-item --request-items file://ContentCatalog.json
aws dynamodb batch-write-item --request-items file://Forum.json
aws dynamodb batch-write-item --request-items file://Post.json
aws dynamodb batch-write-item --request-items file://Comment.json
What does this command do?
The aws dynamodb batch-write-item command is used to load or insert multiple items into DynamoDB tables!
--request-items tells DynamoDB that the items are currently stored inside a file that it'll need to retrieve from.
file:// then tells DynamoDB that the file is stored locally in the CloudShell environment, with the name FILENAME.json.
💡 How does DynamoDB know which table to store which data? Each .json file you upload tells DynamoDB which table the items should go to!
- After each data load, you should get this comment saying that there were no Unprocessed Items.
What's an Unprocessed Item?
Unprocessed items are records that weren't written to your database! If you see an unprocessed item, an error happened while loading your data into DynamoDB.
Did you run into an error with any of your files?
Here's what you can do:
- Find the name of the file that has an error e.g. for this screenshot, it's ContentCatalog.json
- Run nano FILENAME. Replace FILENAME with your file e.g. nano ContentCatalog.json
- Paste the JSON code for your file - click here to get a folder with all datasets.
- This is what your terminal should look like:
- Press Ctrl + X on your keyboard, then press Enter on your keyboard to finish editing ContentCatalog.json.
- Run aws dynamodb batch-write-item --request-items file://ContentCatalog.json again.
- Ask the NextWork community if you're still stuck!
View and update your loaded data
- Head back to the DynamoDB console.
- Select Tables from the left hand navigation panel.
- Pick the ContentCatalog table.
- Select Explore table items on the top right.
- Wooooohoo! Your items are now on display.
What am I seeing?
The data you loaded when you ran aws dynamodb batch-write-item is here!
We can see now that the table has a partition key of Id (the very first column!), and there are 6 items in the table.
💡 What is the partition key again? Partition keys are like tags that DynamoDB uses to organise the table's data. When you search for an item in your table, DynamoDB will need its partition key!
- Scroll through all the columns, can you tell what this data is showing?
- Take a look at the ContentType column. Some items are Projects and some items are Videos.
- This means the ContentCatalog Table stores NextWork's entire collection of content, which includes step-by-step projects (Projects) and a wide range of videos (Videos).
- Click into a Project e.g. click into the item with the Id 1.
- Wow! You get to see all the attributes in this Project right away.
- When you're done, click Save and close.
- Try opening a Video item e.g. the item with Id 203.
- Ooo this item has its own list of attributes...
- Are all the attributes in the previous item here?
- Nope! For example, we don't see Difficulty or ProjectCategory here.
- This is the reason why DynamoDB is known for its flexibility - every single item can have their own set of attributes. Just because StudentsComplete is an attribute in the first item (with Id 1), that doesn't mean there's a StudentsComplete attribute in the other items in this table.
What's the difference between this and a relational database?
Relational databases would need each row to have the same number of columns. So if you added StudentsComplete as a new column in a relational database, every item in that database would need to have a StudentsComplete value too, even if it doesn't apply.
This has huge impacts on a DynamoDB vs a relational database's flexibility and speed!
- Flexibility - every item having their own unique set of attributes is a huge advantage when items in a table could look different from each other. For example, e-commerce sites and shopping carts need to store different types of products with different attributes in the same place.
- Speed - DynamoDB tables can use partition keys to split up a table and quickly find the items they're looking for. Relational databases have to scan through the entire table to find data, which can slow down performance.
💡 When would someone pick relational databases over non-relational? Relational databases use SQL, which makes handling complex queries a lot more straightforward!
The strictness of a relational database's schema also means data is kept precise, accurate and consistent, which can be helpful for situations where the quality of data is a top priority (e.g. healthcare systems often opt for relational databases to keep patient records).
Run your first query with the console
Woohoo! We've got our DynamoDB tables AND loaded in data. Nice work!
Now let's extract some insights from our data using queries.
In this step, get ready to:
- Run basic queries on your DynamoDB tables.
- Find data using partition and sort keys.
- Stay on the ContentCatalog table, but this time let's select Query under the heading Scan or query items.
- Let's enter an Id (Partition key) to query for a specific item. Enter 201
What is a Partition key?
Think of a partition key as the filter that DynamoDB will use to split up and find data.
Partition key values don't have to be unique. For example if "Color" is a partition key, items can share partition key values like "Blue", "Green", "Red" and more.
In this table, each item does have a unique partition key value, which helps with finding a single, specific item even faster
- Select Run.
- Oooo, now only the specific item you've searched for is in your table.
- Are you ready to query another table?
- From the Tables section on the left hand side of the console, select Comment.
- Expand the Scan or query items arrow.
- Select Query.
- Ooo, notice anything different?
- There's Id (Partition key) like last time, but there's also a new Sort key called CommentDataTime underneath.
What is a Sort key?
A sort key is a secondary key used to filter your query results again! Sort keys work after the partition key i.e. you still have to use the partition key to split up your data first, and then the sort key partitions your data again.
Sort keys are optional, which is why our ContentCatalog table could still work without one!
Extra for Experts: A sort key is a great (and oftentimes, necessary) tool for tables where items can share a partition key value! In those situations, the partition key + sort key should form a unique combination. That combination is called our primary key so you can still query for single, unique items in that table.
- Before we run any queries, check out the items in Comment.
What is this table storing?
This table is storing comments left on posts in the NextWork community.
It might help to check out the Posts table too - notice that there are two items there, and each item is a post that was made in the community.
A Comment is a comment made on that post, which is why their ID is the name of the original post!
- Now try to query this so that you're only seeing comments to the post I have a question/Just Complete Project #7 Dependencies and CodeArtifacts. Only show comments that were posted from the 1st of September, 2024.
- Under Id (Partition key), enter I have a question/Just Complete Project #7 Dependencies and CodeArtifacts
- Under CommentDateTime (Sort key), switch the Equal to dropdown to Greater than.
- Enter 2024-09-01 as your sort key.
- Select Run.
- Bingo!
- Here's your next challenge: how would you look for all comments posted by User Abdulrahman?
- Clear the Id (Partition key) and CommentDateTime (Sort key).
- Expand the Filters arrow.
- Enter PostedBy as the Attribute name.
- Enter User Abdulrahmanas the Value.
- Select Run.
There's an error!
That's right - you have to use the Id (Partition key) when you query items.
This teaches us two things:
- Data modelling is sooo important - data engineers think very carefully about what queries they should prepare and design for before they upload any data!
- We probably wouldn't have this kind of issue if we could use SQL, so this is a situation where it would be beneficial to use a relational database!
💡 What's data modelling? Data modeling in DynamoDB is all about planning how to set up your tables, e.g. what should be a table's partition keys and sort keys?
This planning stage is essential because it affects how easily you can get to your data and how quickly your database responds later on. If you don’t get this part right, like if you use the wrong keys in your queries, finding the data you need can become really slow or even impossible (like our scenario here!).
Run Queries with AWS CLI
Let's see how you could run the same queries in the CLI!
In this step, get ready to:
- Query your DynamoDB tables with AWS CLI commands.
- Head back to your AWS CloudShell environment, this time running this AWS CLI command:
aws dynamodb get-item \
--table-name ContentCatalog \
--key '{"Id":{"N":"201"}}'
What does this command do?
aws dynamodb get-item is the AWS CLI command to get a single item from a DynamoDB table. --table-name ContentCatalog tells DynamoDB that the item belongs in the table called ContentCatalog. --key '{"Id":{"N":"201"}}' tells DynamoDB that the item has the ID "201". The {"N":} means the ID is a number (N for Number).
- Next, run this command.
aws dynamodb get-item \
--table-name ContentCatalog \
--key '{"Id":{"N":"101"}}' \
--consistent-read \
--projection-expression "Title, ContentType, Services" \
--return-consumed-capacity TOTAL
- Nice - here's your output!
What does this command do?
This command is just like the get-item command we ran before, but this time with some extra spice 🌶️ We're adding extra query options that tells DynamoDB exactly how we want it to get the item:
- --consistent-read : you want a strongly consistent read.
- --projection-expression: you only want to know some of the item's attributes.
- --return-consumed-capacity : you want to know how much capacity was consumed by the request.
💡 What is a strongly consistent read? A strongly consistent read means DynamoDB will give you the guaranteed most recent version of that item. By default, a read from DynamoDB uses eventual consistency, which is a lower cost option that is faster and retrieves a version of the data that might not be the most updated version.
Eventual consistency is no issue if you're using a small dataset (consistency is usually reached within a second of updating the data anyway). But, strongly consistent reads could be your choice for situations where your app needs the most recent data to make important decisions e.g. financial trading apps that need the most recent market data.
- Before we move on, check your output again - did it return any item at all?
- Look closely at the command you ran... is there any item in ContentCatalog with the Id 101? (Nope!) 😉
- Run the command again, this time with a valid Id and back to an eventually consistent read:
aws dynamodb get-item \
--table-name ContentCatalog \
--key '{"Id":{"N":"202"}}' \
--projection-expression "Title, ContentType, Services" \
--return-consumed-capacity TOTAL
- We will see that eventually consistent reads consume half as much capacity - which is why it's the default for DynamoDB read operations.
Set up a transaction
Amazing work 🤩
For our grand finale, you might notice that our tables contain lots of related data!
- Each Forum contains multiple Posts.
- Each Post contains multiple Comments.
How does that impact how we manage our data? Let's find out...
In this step, get ready to:
- Handle related data across DynamoDB tables.
- Update two tables in a single AWS CLI command.
- Check out the Forum table - there's a count for the number of Posts, and a count for the number of comments i.e. comments in each forum.
- That means if we add a new Comment to the Comment table, the database needs to increase the Comments count in the related Forum item.
- How would you make sure that when a students posts a new Comment, the number of comments recorded in a forum also updates?
- Let's try do this using AWS CLI.
- Run this command in your AWS CloudShell environment:
aws dynamodb transact-write-items --client-request-token TRANSACTION1 --transact-items '[
{
"Put": {
"TableName" : "Comment",
"Item" : {
"Id" : {"S": "Events/Do a Project Together - NextWork Study Session"},
"CommentDateTime" : {"S": "2024-9-27T17:47:30Z"},
"Comment" : {"S": "Excited to attend!"},
"PostedBy" : {"S": "User Connor"}
}
}
},
{
"Update": {
"TableName" : "Forum",
"Key" : {"Name" : {"S": "Events"}},
"UpdateExpression": "ADD Comments :inc",
"ExpressionAttributeValues" : { ":inc": {"N" : "1"} }
}
}
]'
What does this command do?
This command is a transaction, which is a group of operations that all have to succeed - if any of the operations in the group fails, none of the changes get applied. This makes sure that any change to your database is consistent across all your tables!
In this transaction, you are recording a new comment made by User Connor!
- The Comment table needs to update - you're adding a new item to the table with all the details of Connor's comment.
- The Forum table also needs to update - Connor's comment was made on a post in the Events forum, so the number of comments in that forum should go up by 1.
💡 So I have to update both tables, not just one? Yup, that's why a transaction is so helpful to make sure the update is done consistently across both tables.
This is an example situation where the AWS CLI is quite beneficial over the AWS Management Console - transactions don't exist in the console, and updating things manually with clicks can definitely lead to mistakes!
- Whoosh! We've updated both the Comment and Forum tables.
- Let's check our work.
- Run this command in CloudShell to view your Events forum, and you'll see the Comments count go up from 0 to 1.
aws dynamodb get-item \
--table-name Forum \
--key '{"Name" : {"S": "Events"}}'
- Here's what you'll see:
- Check the number of comments in the Events table in your console now too:
Nice work!
You've just learnt how to create a DynamoDB table AND create some fantastic queries.
Thanks for sharing all those juicy insights! You're off to a rocking start as NextWork's data engineer ⭐️
Delete Your Resources
Delete Your Resources
Before diving into the steps for deleting your resources, why not challenge yourself to delete everything in this project on your own?
Keeping track of your resources, and deleting them at the end, is absolutely a skill that will help you reduce waste in your account.
Important
Deleting resources that are not actively being used stops you getting charged and is a best practice. Not deleting your resources will result in charges to your account.
- [] Delete the DynamoDB tables
- Try deleting your tables using AWS CloudShell. This command will delete all of their items too.
aws dynamodb delete-table --table-name Comment
aws dynamodb delete-table --table-name Forum
aws dynamodb delete-table --table-name ContentCatalog
aws dynamodb delete-table --table-name Post
Note
Does your terminal look like this after running a delete command?
Sometimes CloudShell enters into a text editor mode, depending on the terminal's response to one of your commands. This is what the text editor mode looks like, and your terminal would not accept regular AWS CLI commands until you exit that mode or editor. Type :q (which means 'quit') and press Enter on your keyboard to release your terminal from the text editor mode.
Do I need to delete AWS CloudShell too?
Nope. The CloudShell environment doesn't cost you to run!
Nice Work!
Nice Work!
ALL DONE!!!!! 🥳
Today you've learnt how to:
- 🌟 Create a DynamoDB table: You could do this with both the AWS Management Console and AWS CLI.
- ⬆️ Load data into DynamoDB: Using AWS CLI, you loaded four files into matching DynamoDB tables. Such a time saver!
- 🔎 Query data in the DynamoDB console: This came with lots of learnings about partition and sort keys, and the benefits (and limits) of using DynamoDB compared to relational databases!
Ready to quiz yourself? You got this! 💪
It's wild that all these learnings are packed in one project. Great work and we'll see you in the next one!
p.s. Does it say "Still tasks to complete!" at the bottom of the screen?
This means you still have screenshots left to upload, or questions left to answer!
- Press Ctrl+F (Windows) or Command+F (Mac) on your keyboard.
- Search for the text Return to later.
- Jump straight to your incomplete tasks!
- 🙋♀️ Still stuck? Ask the community!