Transcribe Audio Files with AI
Learn how to use Amazon Transcribe to transcribe audio files with AI!
Introduction
β‘οΈ 30 second Summary
Welcome to this AWS x AI project on Amazon Transcribe π¬
Audio data is everywhere, from podcasts and lectures to customer service calls and voice notes.
Being able to automatically transcribe this data opens up a world of possibilities - imagine turning every audio file into searchable text, subtitles or blogs!
In this project, you'll gain hands-on experience with transcribing videos and audios using AWS services.
This project is ideal for roles like Data Analysts, AI Developers, and Cloud Architects, where it's important to use AI tools like Transcribe to develop app features (e.g. voice commands) or improve accessibility.
Let's get ready to...
- Store a video with Amazon S3.
- Transcribe your video with Amazon Transcribe.
- Write custom vocabularies to improve your transcription's accuracy.
- Use custom filters to automatically remove unwanted words.
- Transcribe your own voice as you talk live!
If you're up for a bit of a challenge, quiz yourself on the key concepts up ahead in this project.
Store a Video File
Let's start by creating a storage space for the video we'd like to transcribe today.
We'll use Amazon S3 (Simple Storage Service), our go-to service for storing all sorts of data in the cloud. Think of it as a massive, secure, and reliable hard drive in the sky.
In this step, you're going to:
- Create an S3 bucket.
- Upload a video to the bucket.
Log into the AWS Management Console
- Log in to the AWS Management Console as your IAM Admin user.
- Make sure you're in the AWS region closest to you.
- Head to the S3 console.
What are S3 buckets?
S3 buckets are like folders in the cloud where you can store your files, just like folders on your computer.
These buckets hold objects. An object is simply a file (in this case, video or audio files).
π‘ Why are we using S3 in this project?
The star AWS service in this project is Amazon Transcribe, which is a service that converts audio and video into text - but it doesn't store the files. You need to store the files somewhere S3 first, so we're using S3. The finished transcription will also need to be stored in an S3 bucket (or another storage service).
Create an S3 Bucket
- Select Create bucket.
- Enter a unique Bucket name, like nextwork-project-transcribe-[your-name]-[random-numbers].
- Replace [your-name] and [random-numbers] with your actual name and a random string of characters.
- Leave all other settings as default.
Why are we leaving all other settings as default?
We would adjust the other settings if we have specific access controls, like making the objects inside publicly accessible.
In this project, we only need Transcribe to have access to our video. No special access controls are needed to make that happen - Transcribe automatically has full access to any S3 bucket with transcribe in the name.
- Select Create bucket.
- Click into your created bucket.
Download Your Video File
- Download the following video file:
- Right click on the link below, and select Save link as...
- 1-min-clip.mp4
- Let's verify what you've downloaded.
- Head to the Downloads folder in your local computer.
- Open 1-min-clip.mp4 in your local computer.
- Do you see a video file?
What's this video?
This is the video that we'll be transcribing today!
We'd recommend taking a minute and skim through the clip, so you know what to expect from the transcription.
P.S. This clip is a 1-minute excerpt of our project demo on containers and ECR.
Upload Your Video File
- Head back to the S3 console.
- Select Upload.
- Select Add files.
- Select 1-min-clip.mp4 in your Downloads folder.
- Select Upload.
- Upload success!
Congrats! That's your project's video file in the bucket now, ready for transcribing.
Run Your First Transcription
In this step, let's run our video through Amazon Transcribe for the first time.
Amazon Transcribe usually makes a few mistakes in its first time transcribing jargon or words that are hard to hear. Get ready to analyze your transcription and spot inaccuracies too!
In this step, you're going to:
- Create your first transcription with Amazon Transcribe.
- Analyze the generated transcription.
Run a Baseline Transcription
- Head to the Amazon Transcribe console.
What is Amazon Transcribe?
Amazon Transcribe is a service that converts speech to text.
It uses machine learning to understand and process human language. This lets it recognize different voices, accents, and understand the context of what is being said. It can transcribe speech from a variety of sources, like videos, voice recordings and live streams.
- Select Transcription jobs from the left hand navigation panel.
- Select Create job.
- Enter a job name, such as nextwork-project-baseline-transcription.
- For Language settings, we'll keep the default setting of Specific language.
What do language settings means?
Language settings tell Transcribe service what language is spoken in your audio file.
- If you know the language, choose Specific language for the best results.
- If you don't know the language, choose Automatic language identification. Amazon Transcribe can identify the dominant language spoken in your media file in the first three seconds of speech
- If there are multiple languages, choose Automatic multiple languages identification. With this option, Transcribe will use AI to detect when another language is being used, and switch to transcribing that language. duce accurate transcriptions.
- For Language, we'll keep the default setting of English.
- For Model type, select General model.
What are model types?
A model type is a set of instructions that tells Amazon Transcribe how to convert speech to text.
Amazon Transcribe uses different models to understand different types of speech. There is a general model for everyday speech, but you can train your own models for conversations with specific jargon and context.
For example, you can set up a model that helps Transcribe better understand medical or legal conversations, so the transcription is more accurate.
π‘ Is the model type related to AI or machine learning?
Absolutely! The model type is key to how Amazon Transcribe uses AI. It represents the machine learning algorithms and data that Transcribe uses to convert spoken language into written text.
These models learn from lots of audio samples to recognize different accents, terminology, and background noise.
Choosing the right model type, like a general one for everyday speech or a custom one for technical subjects, shows Transcribe's AI power. Transcribe can personalize its transcription to your needs and improve how accurately it converts speech to text.
π‘ Extra for Experts: How can I create my own model type?
Creating your own model type happens right inside the Transcribe console too. You'd need to upload training data, such as 10,000+ words of accurate and relevant transcripts. Then, Transcribe would take 2-3 hours to train a custom model, taking most of the training heavy lifting from you.
- For Input data, this is where we tell Transcribe the file we want to process!
- Select Browse S3.
- Find the bucket that contains your uploaded audio file. Its name should start with nextwork-project-transcribe-
- Click into your bucket.
- Select 1-min-clip.mp4 and select Choose.
- For Output data, keep the default option.
What is output data?
Output data is where Transcribe saves the results of your transcription job.
You can choose to save your transcriptions in a default S3 bucket (i.e. Transcribe creates and manages one for you) or your own bucket.
- Select Next.
- Skip the additional settings - we'll cover each of them later in this project! For now, we'll generate this transcription without extra settings.
Why are we not using any of the extra settings?
When you're using a transcription service, it's industry practice to do the first run without using any custom settings. This is called a baseline transcription.
It's like taking a "before" picture before trying to improve your transcriptions. You want to see how the service performs with the raw audio before you start making any changes.
Your baseline transcription gives you a starting point to compare to later when you add custom settings.
- Select Create job.
- Success!
- Wait 1-2 minutes for the transcription job status to change from In progress to Completed. While we wait...
- Check whether the transcription job is now complete.
Review the Baseline Transcription Output
- Once completed, select the job. Let's view the transcription!
- Scroll down and take a look at the Transcription preview.
- Are there are any inaccuracies or ways this transcription can improve?
How can this transcription improve?
- repositories is misspelled as repositoriesies
- Speech fillers like um can be excluded for a better reading experience.
- 403 Forbidden, a technical term, is mis-transcribed to 4 or 3 forbidden
- player A is not an inaccuracy, but keeping in mind the context of this video, it should be capitalized to Player A as it's a title.
π‘ Why do inaccuracies exist in a transcription?
Transcription errors can happen for a bunch of reasons, like background noise, someone speaking unclearly, or just the tricky nature of language itself.
Even though tools like Amazon Transcribe are pretty smart, they can still get tripped up by accents, local slang, or special terms that aren't in their usual training setup too.
This might lead to some mistakes or missing parts in what gets written down. Also, simple typos can pop up if the model mishears words or bumps into terms it doesn't recognize often.
Nice work creating your very first transcription! Looks like we have a few inaccuracies we'd like to fix in this transcription. Let's start to resolve them now.
Create and Apply a Custom Vocabulary
To start fixing inaccuracies, let's create a custom vocabulary. A custom vocabulary helps Transcribe recognize specific terms that might not be in everyday speech.
This is especially useful for industry-specific jargon (e.g. 403 Forbidden), product names (e.g. Amazon EC2), or unique phrases.
In this step, you're going to:
- Create a custom vocabulary file.
- Upload the file to your S3 bucket.
- Create a custom vocabulary in Amazon Transcribe.
Create the Vocabulary
- In the Amazon Transcribe console, navigate to Customised vocabulary in the left-hand navigation pane.
- Under Manage vocabularies, select Create vocabulary.
- Enter a name for your vocabulary, such as nextwork-project-vocab.
- Select English, US (en-US) as the language.
- Under Create and import vocabulary, select Create vocabulary on the console.
What is this setting for?
To create a new custom vocabulary, you have three options in Amazon Transcribe:
- Upload a custom vocabulary file from your computer. This is great for situations where you already have a prepared list of terms on your local machine e.g it's been handed to you by another engineer.
- Import a custom vocabulary file from an S3 bucket. This is great for users who are managing large datasets or need to integrate with other AWS services. Using S3 makes sure your files are securely stored and easily accessible within the AWS ecosystem.
- Create a custom vocabulary from scratch in the AWS Management Console. This is great for temporary projects and small custom vocabularies.
- Let's tell Transcribe the words we'd like it to know!
What are all these columns in the table?
- The Phrase column is where you enter the specific words or phrases you want Transcribe to recognize.
- The DisplayAs column is where you tell Transcribe how these phrases should appear in the transcription, so they're displayed in a standardized or preferred format.
π‘ Extra for Experts: There are two other columns in the table, what about those?
The SoundsLike and IPA columns are no longer supported for Custom Vocabulary. Things you write in those columns will get ignored, and AWS will take away these columns in the future. SoundsLike and IPA were used for giving phonetic or pronunciation guidance for difficult or unique words.
Tip: You can read more about this in AWS's official documentation.
- Under View and edit vocabulary, select Add row.
- Under Phrase, we'll write player.
- Select the checkbox to save your input.
- Under the DisplayAs column, we'll write Player
- Select the checkbox to save your input.
What will happen to the transcript now?
The next time Transcribe processes your video, it will automatically capitalise player and turn it to Player in the transcription.
Don't worry, you haven't trained Transcribe to do this in all the transcription jobs you run.
Transcribe will only do this in transcription jobs where it's using this custom vocabulary. This means you still have the option to transcribe player without capitalisation in other jobs.
- Select Add row again.
- This time, the Phrase is repositoriesies. Yikes, let's correct that typo!
- Under DisplayAs, enter repositories
- Select Add row again.
- This time, the Phrase is four-or-three-forbidden
- Under DisplayAs, we'll correct the typo to 403 Forbidden
- Are there any other typos or updates you can think of? Add them to the vocabulary!
I ran into an error
Watch out - Transcribe is quite strict on the characters it can allow in the Phrase column:
- Make sure you're not using any numbers i.e. 0-9
- Make sure there are no spaces in the Phrase column.
You can check out the full guide on writing phrases here. If you're stuck, ask the NextWork community!
- Select Create vocabulary.
- Off we goooo, your custom vocabulary should be in the pending state now.
- Wait for the vocabulary status to change from Pending to Complete. This might take a few minutes, so while we wait...
Recap: Why did we create a custom vocabulary?
A custom vocabulary file is a list of words or phrases that you want Amazon Transcribe to recognize.
It helps Amazon Transcribe understand words that it might not normally recognize, especially technical terms ("403 Forbidden") or acronyms ("Amazon EC2").
This makes the transcription more accurate, especially for audio that uses specialized language.
- Check that the vocabulary's status now says Ready.
Nice work creating your custom vocabulary!
This should help solve the inaccuracies in our transcription with '403 Forbidden' and 'player'.
Hmmm, we still have one more thing to solve though... we still need to address the filler words in our transcription e.g. ("um"). Luckily, there's a tool for this in Transcribe too.
Create and Apply a Vocabulary Filter
Next, let's create a vocabulary filter in Transcribe.
Filters remove unwanted words or phrases from our transcripts, like filler words or sensitive information. This will make our transcripts cleaner and more professional.
In this step, you're going to:
- Create a vocabulary filter file.
- Create a vocabulary filter in Amazon Transcribe.
Create the Vocabulary Filter
- In the Amazon Transcribe console, navigate to Vocabulary filtering in the left-hand navigation pane.
- Select Create a vocabulary filter.
What's the difference between a custom vocabulary and vocabulary filters?
Recap: A custom vocabulary helps Amazon Transcribe understand specific words or phrases that it might not normally pick up.
On the other hand, vocabulary filters block out words you donβt want to appear in your transcripts. This could be swearing, filler words like "um" or "uh," or sensitive information you prefer to keep private.
π‘ Extra for Experts: Can't I use custom vocabularies to mask or remove unwanted words?
Technically, you could use custom vocabularies to display certain words as blank spaces or asterisks... but you'd have to define that indvidually for every single unwanted word.
A vocabulary filter is a much more efficient tool for this, because you'd only have to define the unwatend words in a list, and it applies a blanket rule across all of them.
- Enter a Name for your filter, such as filler-words-filter
- Select English as your Language.
- Under Vocabulary input source, select File upload.
What is an input source?
An input source is where your vocabulary filter comes from. The filter is usually imported as a text file, either uploaded directly from your computer or as a file stored in an S3 bucket.
Aha - as you might imagine, we don't quite have a filter file to upload yet... let's do that now.
Create the Vocabulary Filter File
Why are we creating a vocabulary filter file?
A vocabulary filter file tells Transcribe which words to remove or censor in the final transcript. This is helpful for removing filler words, sensitive information, or anything else you don't want in your transcript.
The instructions for creating a filter file depends on your computer's operating system
Pick the one that works for you:
π MacOS
- In your local computer, open TextEdit.
- Select New Document.
- At the top menu bar, select Format.
- Select Make Plain Text.
I saw Make Rich Text instead
Too good! That means your file is already in plain text. You won't need to select anything π
- Add filler words you want to filter out, with a comma between each. For example:
uh, um, like
- Add other filler words you can think of!
- Save the file by pressing Command + S on your keyboard.
- Title your file filler-words-filter.txt
- Select Save.
- Skip the πͺ Windows and π§ Linux instructions below, and head straight to the heading β¬οΈ Upload Your Filter File
πͺ Windows
- On your computer, open Notepad.
- Click on File in the top menu bar.
- Choose New from the dropdown menu to open a new document.
- Ensure the document is set to plain text by clicking Format on the menu bar and selecting Plain Text if available.
- Type the filler words you wish to filter out, separating each with a comma. For example:
uh, um, like
- Feel free to add any additional filler words as needed.
- Save the document by pressing Ctrl + S on your keyboard.
- Name your file filler-words-filter.txt.
- Click Save.
- Skip the π§ Linux instructions below, and go directly to the heading β¬οΈ Upload Your Filter File
π§ Linux
- Open your preferred text editor (e.g., Gedit, Nano, Vim).
- If using Gedit, simply start a new document. If using Nano or Vim, you might need to start them by typing nano or vim followed by filler-words-filter.txt in the terminal.
- Ensure your document is in plain text mode.
- Enter the filler words you want to exclude, each separated by a comma. Example:
uh, um, like
- Add any additional filler words that come to mind.
- Save the file:
- For Gedit, click Save in the top menu or press Ctrl + S.
- For Nano, press Ctrl + O, then Enter, and Ctrl + X to exit.
- For Vim, type :wq and press Enter.
- Name your file filler-words-filter.txt before saving.
β¬οΈ Upload Your Filter File
- Head back into your Transcribe console and your vocabulary filtering setup page.
- Now that you have your text file ready, we can upload it!
- Find the Vocabulary input source setting.
- Select Choose file.
- Select your filler-words.txt file.
- Head to the bottom of your set up page.
- Select Create a vocabulary filter.
- Great success!
Would you look at that! Nice work going above and beyond with Transcribe by customising your own filters too. This will be so helpful with removing unwanted words in speech.
Run an Enhanced Transcription
Let's put our new custom vocabulary and filter to the test. It's time to run a new transcription job!
We'll also learn some extra features in Transcribe - like identifying speakers and generating subtitles - along the way.
In this step, you're going to:
- Create a new transcription job with your filter and custom vocabulary.
- Review the improved transcription output and subtitles.
Create the Enhanced Transcription Job
- In the Transcribe console's left hand navigation panel, select Transcription jobs again.
- Select Create job.
- Enter a job name, such as nextwork-project-enhanced-transcription
- Under Input data, select Browse S3.
- Find the bucket that contains your video. Its name should start with nextwork-project-transcribe-
- Click into your bucket.
- Select 1-min-clip.mp4 and select Choose.
- For Output data, keep the default option.
- Select SRT (SubRip) as a Subtitle file format.
Why are we choosing SRT as the output format?
Subtitles are text files that display words spoken in a video. They help people who are deaf, hard of hearing, or speak different languages to understand a video or audio.
Because of this, many apps that serve any video or audio content, from YouTube to internal company portals or games, could use subtitles.
Amazon Transcribe can generate subtitles in two formats: WebVTT (.vtt) and SubRip (.srt) to work with different types of media players and editing tools. SubRip is the more popular format that works with most video players and editing software, so we'll go with that.
- Select Next.
- Under Audio settings, enable Audio identification.
- Check Speaker partitioning and set the maximum number of speakers to 2
Why is speaker partitioning?
Speaker partitioning is a feature that helps you label different speakers in an audio recording i.e. who said what.
This is helpful for understanding conversations or recordings with multiple people. If you ever want to focus on what a specific person talked about throughout a recording, you can use the partitioned data to filter out everyone else!
In our case, since we have two different speakers in our test video, let's enable this feature and see how Transcribe partitions the transcript.
P.S. speaker partitioning is also called speaker diarization in other transcription services.
π‘ Extra for Experts: Why does the other setting (channel identification) mean?
In more complex audio setups, a recording might combine individual audio tracks. For example, a live band recording with different microphones for different instruments. If you have multiple audio tracks, each track is called a channel.
Transcribe has the ability to identify different channels in a recording, which can make it easier for you to match speech to the correct person. Since this demo video has a relatively simple audio setup, we won't need this setting.
We'll skip the other audio settings.
Extra for Experts: What are alternative results?
Alternative results are different versions of the same audio transcription, each version with a different level of confidence. This is helpful when accuracy is important. You can choose the version that best suits your needs! If you don't have this enabled, Transcribe simply gives you the transcription with the highest confidence score.
π‘ Extra for Experts: What is PHI identification?
Since the language we picked is English (US), Amazon Transcribe has a special setting for complying with a US regulation (called HIPAA). This regulation says information that can be used to identify a patient is protected health information i.e. PHI.
Amazon Transcribe can use AI to flag PHI in a transcript, and even label the type of PHI e.g. financial, health, geographic, biometric information.
By knowing where PHI exists in transcripts, companies can apply stricter security tools like encryption, access controls, and redaction, to those parts of the data. It can also secure patient information even if a data breach were to happen.
π‘ Extra for Experts: What is toxicity detection?
Toxicity detection is a way to find harmful or inappropriate content in speech.
It works by analyzing the words, tone of voice, and emotions in speech, and assigning a score based on how toxic it is. If the toxicity score is high, the speech is flagged as toxic and categorized too.
- Under Content removal, enable Vocabulary filtering.
- Under Filter selection, select filter-words-filter
- We'll keep the Vocabulary filtering method as Mask vocabulary.
Why are we picking Mask vocabulary?
Mask vocabulary replaces unwanted words with asterisks (***) to protect privacy or keep content appropriate. This way, the word is still there, but it's not readable.
In this project, masking makes it very obvious whether Transcribe has successfully identified any of our filler words - if it did, they'll show as ***. This makes analysis easy.
- Finally, under Customisation, enable Customised vocabulary.
- Under Vocabulary selection, select nextwork-project-vocab.
- Select Create job.
- Yipee! Look at it go.
Exciting! You've just run a new transcription job that uses your custom vocabulary and filters.
Review Your Enhanced Transcription
What do you think your new transcription will look like?
Let's find out in the grand finale... it's time to review your newest transcription!
- Wait for the transcription job status to change to Completed.
- Once completed, select the job. Let's view its transcription.
- Scroll down to Transcription preview.
- Do you see any difference between this transcription and the baseline? Note the improvements.
403 Forbidden is still wrong in this transcript
Oh no! If Transcribe is still making the same mistake, it means your setup still needs to be adjusted for better accuracy.
Engineers typically improve their setup by adding more variations to their custom vocabulary, or training a custom model on a broader set of data that captures their use case better. See if you can work out how to do it!
For now, the learning here is that customization doesn't always give us the solution we want. The results get better with more data and model training.
- Next, select the Audio identification tab.
- Oh nice! You can see Speaker 0 and Speaker 1 clearly identified now.
What is Audio identification useful for?
As a recap, audio identification is a technology that can recognize and tell the difference between different sounds and voices.
It's used in...
- Meeting or podcast transcripts to identify different speakers.
- Security systems to detect suspicious noises.
- Customer support to filter for what customers are saying
- Apps that need to react to specific sounds.
- Next, select the Subtitles tab.
- Awesome. This time, we can see subtitles in SRT format popup.
- Select the Download button at the top right corner.
- Now you'll find your subtitles as an export option.
What are subtitles useful for?
As you might imagine, subtitles make videos and entertainment accessible to more people. They help people who are deaf, hard of hearing, or speak different languages understand a video or audio.
Apps that serve any video or audio content, from YouTube and Netflix, to internal company portals or games, use subtitles.
Great work learning how to transcribe a video with Transcribe!
You've also created a custom vocabulary and filter to improve your transcription's quality, and the results were awesome.
Secret mission
Welcome to your π€« exclusive π€« secret mission!
Your mission, should you choose to accept it, is to use Transcribe for real-time transcription.
π In this secret mission, get ready to:
- Start a real-time transcription stream.
- Speak into your microphone and see Transcribe generate a transcription instantly.
- Showcase more advanced knowledge, like the use cases of real-time transcription, in your project documentation. Stand out from the rest!
Real-Time Transcription
Delete your resources
Delete your resources
Now that we've explored the power of Amazon Transcribe, it's time to clean up the resources we created. This is important to avoid incurring unnecessary costs. It's like tidying up your kitchen after cooking - you don't want to leave a mess behind.
Resources to delete:
- Delete the transcription jobs.
- Delete the custom vocabulary.
- Delete the vocabulary filter.
- Delete the files from your S3 bucket.
- Delete the S3 bucket itself.
β STOP
Before diving into the steps for deleting your resources, why not challenge yourself to delete everything in this project on your own?
Keeping track of your resources, and deleting them at the end, is absolutely a skill that will help you reduce waste in your account.
π STEPS BELOW:
- Delete Transcription Jobs
- In the Amazon Transcribe console, navigate to Transcription jobs from the left hand navigation panel.
- Select nextwork-project-baseline-transcription
- Select Delete.
- Select Delete again.
- Select nextwork-project-enhanced-transcription
- Select Delete.
- Select Delete again.
- Delete Custom Vocabulary
- Navigate to Customised vocabulary from the left hand navigation panel.
- Select your custom vocabulary nextwork-project-vocab
- Select the Actions dropdown.
- Select Delete.
- Select Delete again.
- Delete Vocabulary Filter
- Navigate to Vocabulary filtering from the left hand navigation panel.
- Select your vocabulary filter filler-words-filter
- Select Delete.
- Select Delete again.
- Delete S3 Files and Bucket
- Head to the S3 console.
- In your S3 bucket, select your bucket. Its name should start with nextwork-project-transcribe
- Select the video file you've uploaded.
- Select Delete.
- At the bottom of the page, type permanently delete
- Select Delete objects.
- Head back to the Buckets page from the left hand navigation panel.
- Select your bucket again. Its name should start with nextwork-project-transcribe
- Select the bucket.
- Select Delete.
- Enter the name of your bucket.
- Select Delete bucket.
That's a wrap!
That's a wrap!
Wow! You've just transcribed a video AND used some flashy techniques to improve your transcription.
You've learned how to:
- πͺ£ Store a video with Amazon S3.
- π£οΈ Transcribe your video with Amazon Transcribe.
- π€ Write custom vocabularies to improve your transcription's accuracy.
- π€¬ Use custom filters to automatically remove unwanted words.
- π Experiment with real-time transcription.
p.s. Does it say "Still tasks to complete!" at the bottom of the screen?
This means you still have screenshots left to upload, or questions left to answer!
- Press Ctrl+F (Windows) or Command+F (Mac) on your keyboard.
- Search for the text Return to later.
- Jump straight to your incomplete tasks!
- πββοΈ Still stuck? Ask the community!