Load Data into DynamoDB
Let's learn about non-relational databases with Amazon DynamoDB!
Introduction
⚡️ 30 second Summary
Welcome to the NextWork team!
You've just joined the NextWork team as our Data Engineer, and we're thrilled to have you on board 👋
As we build the NextWork community, we're wanting to store our projects, videos and community activity in a database.
Do you think you can set up DynamoDB to house all our data?
In this project, get ready to...
- 🌟 Create a table with Amazon DynamoDB
- ⬆️ Upload data into DynamoDB using AWS CloudShell
- 🔧 Edit data in your DynamoDB tables
Want a complete demo of how to do this project, from start to finish? Check out our 🎬 walkthrough with Natasha 🎬
If you're up for a bit of a challenge, quiz yourself on the key concepts up ahead in this project.
Before we start Step #1...
Before we get started, it's important that you know what we're trying to do today.
Login with your IAM user
For this project you'll need your IAM user, not your root user.
First things first... do you have an IAM user?
No
Oooo it's the start of a new era!
If you don't have an IAM user yet - here are the steps to create one (this takes less than 10 mins).
What is an IAM user? Why are we setting one up?
In AWS, a user is a person or a computer that can do things on the AWS cloud.
When you create an AWS account for the first time, the login you get is called the root user of the AWS account. AWS actually recommends to not use your root user for everyday tasks to protect it from security breaches.
You should create IAM users instead. If a root user is a master key to your AWS account, think of IAM users as key copies. IAM users have separate usernames and passwords to your root user, and you can set them to have limited access to your account's resources.
- Head to your AWS Account as the root user.
- Open the AWS IAM console.
- From the left hand navigation panel, choose Users.
- Choose Create user.
- For the User name, name it:
[[YOURNAME="enter your name"]]-IAM-Admin
- Make sure to select the checkbox next to Provide user access to the AWS Management Console - optional.
Note
This does not apply to all accounts, but if you're prompted with a pop up panel that says Are you providing access to a person?, choose I want to create an IAM user.
- For the console password, choose Custom password.
- Type in a password that you will be able to remember/access in the future.
Top tip
You will use this password for all future projects, so make sure to choose a secure one!
- Deselect the checkbox for Users must create a new password at next sign-in - Recommended.
- Choose Next.
- In the permissions set up page, choose Attach policies directly.
- From the list of Permissions policies, select AdministratorAccess.
- Choose Next.
- Choose Create user.
- Voilà - you've just created your new user! Stay on this page.
- Choose Download .csv file.
- Copy the Console sign-in URL.
- Now you're ready to start using your IAM user. 🏁
- Log out of your root user's AWS Account.
- Paste and go to your copied console sign-in URL.
- Open your downloaded .csv file containing your user's access instructions.
- Log in using your IAM user's username and password in the .csv file.
- Once you're logged in, you're ready to use your IAM user for this project! Make sure to keep the login details safe - you'll need them for the entire 6 Day DevOps Challenge!
Yes
Nice! Log in to the AWS Management Console with your IAM Admin User.
Note
PLEASE make sure you log in to your IAM Admin User instead of the root user - it's truly best practice for account security.
Create your first DynamoDB table
Let's start with our most fundamental ingredient... a DynamoDB table! Once we have this, we can learn how to populate it with data (and then query it for insights in the next project). Lovely!
In this step, get ready to:
- Create a DynamoDB table from scratch.
- Head to the DynamoDB console in your AWS Management Console.
- From the left hand navigation panel, select Tables.
- Select Create table.
What is DynamoDB?
Amazon DynamoDB is a non-relational database service.
Non-relational databases use structures other than rows and columns to organise data.
Extra for Experts: DynamoDB can also be described as a NoSQL database, or a key-value database. A NoSQL database means you would not use SQL to query it, while key-value is a specific way to store data that's flexible and efficient.
💡 What is a DynamoDB table? In Amazon DynamoDB, all data is organized into tables! Unlike relational databases which use rows and columns, DynamoDB tables use items and attributes. Let's see what these items and attributes look like in the next few steps.
💡 So a DynamoDB table is a table... without rows and columns? How is that possible? It might sound impossible to have a table that doesn't use rows and columns, but it's true that DynamoDB tables aren't like the traditional structure!
You'll see a DynamoDB table in action soon, but in short, imagine if you had a table where each row had a different number of fields, and every single cell can have a different column header. Instead of the typical relational database structure (where each row has the same columns and column headers), database tables are a lot more flexible.
- Give your table a Table name: NextWorkStudents
- For the Partition key, we'll use StudentName
What is a Partition key?
Think of a partition key as the filter that DynamoDB will use to split up and find data.
Partition key values don't have to be unique. For example if "Color" is a partition key, items can share partition key values like "Blue", "Green", "Red" and more.
Later on, partition keys are used to efficiently find and grab items you're looking for (e.g. "find all items that have the Color "Green")! That's why every item in a DynamoDB table must have a value for the table's partition key - otherwise, DynamoDB would have no way of finding that item.
- Next, under Table settings, select Customize settings.
Why can't we use the default settings?
To make sure we're keeping with in the AWS Free Tier! We're going to set up this table with minimal performance requirements, so we don't go beyond the Free Tier's limits.
- Expand the Capacity calculator section.
- Oooo, here's your tip on how AWS charges for DynamoDB.
- Note: you're about to see a cost estimate, but this project and the tables you create are free!
Why is there a cost estimate?
AWS cost estimates always tell you how much a resource might cost you money based on your current settings, but this estimate doesn't consider the AWS Free Tier.
Right before AWS sends you a bill, AWS will check your actual use and won't charge you if you haven't gone over Free Tier limits.
💡 What is item read/second? When we say "1 item read/second" in DynamoDB, it means your application can retrieve one item of data from your database every second.
Changing this number would change your database's performance i.e. larger datasets with lots of items could want hundreds of reads each second!
💡 What is item write/second? Just like item reads, 1 item write/second means your app can save/update one item a second.
- We won't change anything in the Capacity Calculator, but that was useful to understand how DynamoDB pricing works!
- Under Read/write capacity settings, let's make sure costs are as low as possible.
- Under Read capacity, turn Auto scaling off.
What is auto scaling?
Auto scaling can automatically adjust your database's performance (i.e. how fast it can return query results) based on real-time demand.
You'll spot auto scaling available in other services like Amazon EC2 too, where they can automatically adjust the number of instances your application needs.
💡 Auto scaling sounds handy! Why are we turning it off? Auto scaling can help engineers reduce costs in other situations, butttttt if it's not monitored, it also has the power to boost your table's processing power (e.g. if your table suddenly needs to update lots of data). If that ever happens, auto scaling can push your table's settings to go over Free Tier limits!
We're better off making our table stick to our current settings, which are already the lowest cost options.
- Change provisioned capacity units to 1.
- Under Write capacity, turn Auto scaling off.
- Change provisioned capacity units to 1.
- Hmm, what are capacity units?
- Back in your Capacity Calculator, notice that there's something called Read capacity units (RCUs) at the bottom of the panel.
- In your calculator, try changing the number of Item read/second to 2
- The number of Read capacity units in the calculator is still 1!
- Now try changing the number of Item read/second to 3 in the calculator.
- Aha, the number of Read capacity units went up to 2!
What are read capacity units?
Think of read capacity units (RCUs) as a way to measure how many engines DynamoDB is using to operate.
1 read capacity unit (RCU) = DynamoDB is using a single engine to run its read operations. This lets DynamoDB perform a max of 2 reads per second.
So if you decide you actually want DynamoDB to run 3 reads per second, 1 engine won't be enough! The capacity calculator updates RCUs to 2, so now there are two engines powering DynamoDB to run read operations.
- Now try increasing Item write/second from 1 to 2
- Ooo, the number of Write capacity units increases to 2 right away!
What are write capacity units?
Write capacity units (WCUs) are just like read capacity units - they give your DynamoDB tables the engines to edit/update/delete data!
1 WCU = 1 item write/second.
Yup, that means it costs more to run write operations than read operations.
AWS charges your account based on how many RCUs and WCUs you've used each month. If your performance needs are really high, your RCUs and WCUs will go up!
💡 So will I be charged for this project? Nope! To make sure you don't get charged for this project, delete all resources once you're done.
The Free Tier for DynamoDB gives you 25GB of data storage, plus 25 Write and 25 Read Capacity Units (WCU, RCU). This is enough to handle 200M requests per month... all for free.
- Scroll down and select Create table.
- Once your table is created, click into your table's name.
- Pause! ✋
- Take a moment to look around before moving onto the next step. Get comfortable with the screen and the different things on it - there are quite a few panels in this page!
- When you're ready, select Explore table items on the top right hand corner.
- Under Actions select Create item.
- Awesome! From the DynamoDB console, you can enter data manually straight away.
That's handy! Is this a feature for all AWS database services?
Nope, DynamoDB is actually quite unique for letting you do this. All other AWS database services, except for Amazon Keyspaces, don't let you enter data directly from the console.
DynamoDb is designed to be extremely beginner and developer-friendly. By having this feature, developers can quickly update their database structure and see how their apps will respond to these changes in real-time.
Extra for Experts: This feature highlights DynamoDB's status as a fully managed AWS service i.e. you can concentrate more on building your applications instead of database management tasks like maintaining the database’s availability, durability, and scalability.
- Let's add our very first NextWork student to the table. Next to StudentName, enter Nikko
- Ooo, turns out we can even add new attributes! Let's select Add new attribute, this time selecting Number.
- For the new Attributes name, we'll call it ProjectsComplete.
- Enter in a value for the number of projects that Nikko has completed! We'll enter 4.
What is an attribute?
In DynamoDB, an attribute is like a piece of data about an item. In this case, our item is Nikko and the attribute is the number of projects Nikko completed.
Each item in DynamoDB can have multiple attributes. But, unlike relational databases where each row in a table must have the same columns, DynamoDB items can have their own unique set of attributes.
- Select Create item.
Nice! Notice a green confirmation banner that tells us we've consumed 0.5 Read capacity units.
What does the green banner mean?
'Read capacity units consumed: 0.5' means DynamoDB used only 0.5 RCUs to add that item.
In case you're nervous that the AWS Free Tier only provides 25 RCUs a month, this doesn't mean you only have 24.5 RCUs left for the rest of this project.
Think of it like you're driving a car, and the RCU is your speedometer telling you how much data it's processing a second. As long as you're not asking your table to speed up and perform lots of read operations in the same second, you'll be okay!
Check your work. Do you see a Table called NextWorkStudents with a student Nikko?
How is DynamoDB a non-relational database? Isn't this rows and columns that I'm looking at?
This DynamoDB table makes it seem like data is organized in rows and columns, but it's actually quite different from traditional relational databases.
Instead of looking at this as a spreadsheet, look at a DynamoDB table as a list of items (i.e. StudentNames, like Nikko), each with their own list of attributes (e.g. ProjectsComplete).
With DynamoDB, your table can become very flexible and you can start adding new and different attributes for every item! It's like having a spreadsheet, except every row can have a different number of columns and different column headers. This level of flexibility is not possible with relational databases.
You'll see this in action in the next few steps 😉
Create DynamoDB tables with AWS CloudShell
You can probably imagine that creating tables and items one by one isn't the fastest way to do things!
For example, what if you had thousands of students and projects to load into your table?
In this step, get ready to:
- Interact with DynamoDB using a new way... introducing AWS CloudShell!
- Run commands inside CloudShell to create DynamoDB tables.
- At the top of your AWS Management Console, select the icon for AWS CloudShell.
What is CloudShell?
AWS CloudShell is shell in your AWS Management Console, which means it's a space for you to run code! The awesome thing about AWS CloudShell is that it already has AWS CLI pre-installed.
💡 What is CLI? AWS CLI (Command Line Interface) is a software that lets you create, delete and update AWS resources with commands instead of clicking through your console.
You usually have to install AWS CLI into your computer to use it, but in our case, CloudShell already has CLI installed for us (thank you CloudShell 🙏).
Extra for Experts: As you advance in your cloud engineering career, you'll find that the AWS CLI often becomes your go-to tool. Engineers use the CLI to automate tasks and manage AWS resources efficiently using scripts, making it essential for managing your cloud environment in an efficient way.
While the AWS Management Console is fantastic for learning and having a visual guide, the CLI provides the speed and versatility that professionals need for complex tasks.
- Wait 30 seconds for your environment to be ready.
- Run these commands to create new tables:
aws dynamodb create-table \
--table-name ContentCatalog \
--attribute-definitions \
AttributeName=Id,AttributeType=N \
--key-schema \
AttributeName=Id,KeyType=HASH \
--provisioned-throughput \
ReadCapacityUnits=1,WriteCapacityUnits=1 \
--query "TableDescription.TableStatus"
aws dynamodb create-table \
--table-name Forum \
--attribute-definitions \
AttributeName=Name,AttributeType=S \
--key-schema \
AttributeName=Name,KeyType=HASH \
--provisioned-throughput \
ReadCapacityUnits=1,WriteCapacityUnits=1 \
--query "TableDescription.TableStatus"
aws dynamodb create-table \
--table-name Post \
--attribute-definitions \
AttributeName=ForumName,AttributeType=S \
AttributeName=Subject,AttributeType=S \
--key-schema \
AttributeName=ForumName,KeyType=HASH \
AttributeName=Subject,KeyType=RANGE \
--provisioned-throughput \
ReadCapacityUnits=1,WriteCapacityUnits=1 \
--query "TableDescription.TableStatus"
aws dynamodb create-table \
--table-name Comment \
--attribute-definitions \
AttributeName=Id,AttributeType=S \
AttributeName=CommentDateTime,AttributeType=S \
--key-schema \
AttributeName=Id,KeyType=HASH \
AttributeName=CommentDateTime,KeyType=RANGE \
--provisioned-throughput \
ReadCapacityUnits=1,WriteCapacityUnits=1 \
--query "TableDescription.TableStatus"
What does this command do?
This script includes commands to create four new tables in AWS DynamoDB, each with specific attributes and settings.
The four tables created are:
- ContentCatalog Table: This table has a numeric attribute called Id.
- Forum Table: This table has a partition key called Name.
- Post Table: This table has a partition key called ForumName and a sort key called Subject. We'll dive into sort keys soon!
- Comment Table: This table has a partition key called Id and a sort key called CommentDateTime.
- ✈️ Off we goooooo!! AWS CloudShell tells AWS CLI to work with DynamoDB and create those tables for you.
- Tip: if your terminal stops updating, you might need to press Enter on your keyboard to run the last command.
If a Safe Paste panel appears, select Paste.
Bonus tip: You're going to see a lot of multi-line code in this project!
Untick the checkbox Ask before pasting multi-line code so this panel doesn't pop up every time.
- Let's confirm those tables were actually created. Run these wait commands, and wait until they all run and end:
aws dynamodb wait table-exists --table-name ContentCatalog
aws dynamodb wait table-exists --table-name Forum
aws dynamodb wait table-exists --table-name Post
aws dynamodb wait table-exists --table-name Comment
What are wait commands?
When you run a wait command, you're telling your terminal to keep waiting until a condition is finally met. In our case, we're saying "keep waiting - don't finish running this command until this table has been created."
Wait commands are helpful for making sure necessary resources have been created before you move on. Otherwise, future commands that depend on your resources would automatically fail!
Extra for Experts: Can you guess what might happen if you ran a wait command before creating the resource itself?
The wait command will keep running and waiting... eventually it'll decide that the resource doesn't exist and fail i.e. stop!
- Head back into your DynamoDB console and select the Tables tab.
- Confirm that you see four new tables!
- Tip: You might need to refresh your page if you don't see them straight away.
Uhhhh I don't see any data inside the new tables, where are they?
That's totally normal!
Initially, DynamoDB tables are created empty. You need to populate them with data either manually through the DynamoDB console, by using the AWS CLI, or by connecting your database with an app that writes data to them.
You've learnt how to populate data manually first, so now we'll try loading in data using the AWS CLI.
Load Data into Your Tables
Now that we've got our DynamoDB tables set up, we can actually start to load data in.
In this step, get ready to:
- Load some data into DynamoDB tables.
- View and update your loaded data.
- Head back into your CloudShell terminal.
- Download and unzip this zip file with data:
curl -O https://storage.googleapis.com/nextwork_course_resources/courses/aws/AWS%20Project%20People%20projects/Project%3A%20Query%20Data%20with%20DynamoDB/nextworksampledata.zip
unzip nextworksampledata.zip
cd nextworksampledata
We're unzipping a file... into CloudShell?
That's right! CloudShell is an environment that can also handle 1GB of storage, so you can save files inside CloudShell.
And here's a fun tip: you could write and store scripts right in CloudShell to automate repetitive tasks, so you won't need to run AWS CLI commands line by line!
Extra for Experts: CloudShell's storage is persistent, meaning your files and data will stay available across sessions as long as you stay within the 1 GB limit.
- run ls to confirm that all files are now inside your CloudShell environment
- Want to see what's inside these files?
- Run cat Forum.json
What does this command do?
The cat command opens up and lets you read files directly in the terminal!
So when you run cat Forum.json, the terminal will show you all the data inside the Forum.json file right on your screen. This is such a quick way to view or verify the contents of a file without having to open it somewhere else.
So what's inside Forum.json?
Let's break it down! Forum.json contains data that's been formatted specifically for loading into DynamoDB:
- "Forum": tells DynamoDB that the data relates to the Forum table.
- "PutRequest": tells DynamoDB to add a new item into Forum table.
- Then the rest of each PutRequest includes the attributes of this new item! Notice how the first item has five attributes (Name, Category, Posts, Comments, Views) but the second item only has three.
- This is a great example of a DynamoDB table's flexibility - every item can have any number of attributes and attribute titles.
- If we were using a relational database... then both items would need to have a value for all the attribute names in the table! This makes your database bigger (and slower).
- Load the data of all four files into DynamoDB using AWS CLI's batch-write-item command:
aws dynamodb batch-write-item --request-items file://ContentCatalog.json
aws dynamodb batch-write-item --request-items file://Forum.json
aws dynamodb batch-write-item --request-items file://Post.json
aws dynamodb batch-write-item --request-items file://Comment.json
What does this command do?
The aws dynamodb batch-write-item command is used to load or insert multiple items into DynamoDB tables!
--request-items tells DynamoDB that the items are currently stored inside a file that it'll need to retrieve from.
file:// then tells DynamoDB that the file is stored locally in the CloudShell environment, with the name FILENAME.json.
💡 How does DynamoDB know which table to store which data? Each .json file you upload tells DynamoDB which table the items should go to!
- After each data load, you should get this comment saying that there were no Unprocessed Items.
What's an Unprocessed Item?
Unprocessed items are records that weren't written to your database! If you see an unprocessed item, an error happened while loading your data into DynamoDB.
Did you run into an error with any of your files?
Here's what you can do:
- Find the name of the file that has an error e.g. for this screenshot, it's ContentCatalog.json
- Run nano FILENAME. Replace FILENAME with your file e.g. nano ContentCatalog.json
- Paste the JSON code for your file - click here to get a folder with all datasets.
- This is what your terminal should look like:
- Press Ctrl + X on your keyboard, then press Enter on your keyboard to finish editing ContentCatalog.json.
- Run aws dynamodb batch-write-item --request-items file://ContentCatalog.json again.
- Ask the NextWork community if you're still stuck!
View and update your loaded data
- Head back to the DynamoDB console.
- Select Tables from the left hand navigation panel.
- Pick the ContentCatalog table.
- Select Explore table items on the top right.
- Wooooohoo! Your items are now on display.
What am I seeing?
The data you loaded when you ran aws dynamodb batch-write-item is here!
We can see now that the table has a partition key of Id (the very first column!), and there are 6 items in the table.
💡 What is the partition key again? Partition keys are like tags that DynamoDB uses to organise the table's data. When you search for an item in your table, DynamoDB will need its partition key!
- Scroll through all the columns, can you tell what this data is showing?
- Take a look at the ContentType column. Some items are Projects and some items are Videos.
- This means the ContentCatalog Table stores NextWork's entire collection of content, which includes step-by-step projects (Projects) and a wide range of videos (Videos).
- Click into a Project e.g. click into the item with the Id 1.
- Wow! You get to see all the attributes in this Project right away.
- Click Add new attribute, select String from the dropdown.
- Name the new attribute StudentsComplete
- For the value, enter Nikko
- When you're done, click Save and close.
- Nice, you've just added a new attribute - StudentsComplete to an item!
- Now take a look at your Table - looks like StudentsComplete is now a new attribute at the top of the table.
- Does that mean StudentsComplete is now an attribute for all the other items in the table?
- Let's find out. Try opening a Video item e.g. the item with Id 203.
- Ooo this item has its own list of attributes...
- Is your new attribute StudentsComplete here?
- Nope, it isn't!
- This is the reason why DynamoDB is known for its flexibility - every single item can have their own set of attributes. Just because StudentsComplete is an attribute in the first item (with Id 1), that doesn't mean there's a StudentsComplete attribute in the other items in this table.
What's the difference between this and a relational database?
Relational databases would need each row to have the same number of columns. So if you added StudentsComplete as a new column in a relational database, every item in that database would need to have a StudentsComplete value too, even if it doesn't apply.
This has huge impacts on a DynamoDB vs a relational database's flexibility and speed!
- Flexibility - every item having their own unique set of attributes is a huge advantage when items in a table could look different from each other. For example, e-commerce sites and shopping carts need to store different types of products with different attributes in the same place.
- Speed - DynamoDB tables can use partition keys to split up a table and quickly find the items they're looking for. Relational databases have to scan through the entire table to find data, which can slow down performance.
💡 When would someone pick relational databases over non-relational? Relational databases use SQL, which makes handling complex queries a lot more straightforward!
The strictness of a relational database's schema also means data is kept precise, accurate and consistent, which can be helpful for situations where the quality of data is a top priority (e.g. healthcare systems often opt for relational databases to keep patient records).
Nice work!
You've just learnt how to create a DynamoDB table AND load data.
Thanks for getting the ball rolling and creating the database! You're off to a rocking start as NextWork's data engineer.
Next up, we'll be running queries on this database to extract some juicy insights ⭐️
Delete Your Resources
Delete Your Resources
Important
Deleting resources that are not actively being used stops you getting charged and is a best practice. Not deleting your resources will result in charges to your account.
👀 Do you have time for another project today?
Yep, let's go!
Note
You don't need to delete your resources if you're doing the next project in this series today.
Get your documentation and head straight to the next project!
Nope, not today.
- Delete the DynamoDB tables
- Try deleting your tables using AWS CloudShell. This command will delete all of their items too.
aws dynamodb delete-table --table-name Comment
aws dynamodb delete-table --table-name Forum
aws dynamodb delete-table --table-name ContentCatalog
aws dynamodb delete-table --table-name Post
aws dynamodb delete-table --table-name NextWorkStudents
Do I need to delete AWS CloudShell too?
Nope. The CloudShell environment doesn't cost you to run!
Nice Work!
Nice Work!
ALL DONE!!!!! 🥳
Today you've learnt how to:
- 🌟 Create a DynamoDB table: You could do this with both the AWS Management Console and AWS CLI.
- ⬆️ Load data into DynamoDB: Using AWS CLI, you loaded four files into matching DynamoDB tables. Such a time saver!
- 🔎 Edit data in your DynamoDB tables This came with lots of learnings about the differences between DynamoDB and relational databases!
Ready to quiz yourself? You got this! 💪
It's wild that all these learnings are packed in one project. Great work and we'll see you in the next one, Query Data with DynamoDB!
p.s. Does it say "Still tasks to complete!" at the bottom of the screen?
This means you still have screenshots left to upload, or questions left to answer!
- Press Ctrl+F (Windows) or Command+F (Mac) on your keyboard.
- Search for the text Return to later.
- Jump straight to your incomplete tasks!
- 🙋♀️ Still stuck? Ask the community!