AWS Transcribe cost: per second of audio transcribed
Transcribe bills per second of audio (rounded up per request), with higher rates for features like speaker identification, medical, and custom models. Long audio and premium features drive the bill. Here is the per-second model.
Quick answer
Transcribe bills per second of audio transcribed (with a per-request minimum), at higher rates for premium variants like Transcribe Medical and features such as custom language models, with tiered discounts at volume and a free tier. Cost scales with audio duration, so transcribing only what you need, and using standard rather than premium tiers where possible, are the levers.
Amazon Transcribe converts speech to text, billed by the duration of audio processed. Its pricing is simple, per second of audio, so the cost is a direct function of how much audio you transcribe and which tier and features you use.
Per second, by tier and feature
| Variant / feature | Cost |
|---|---|
| Standard transcription | Base per-second rate |
| Transcribe Medical | Higher per-second rate |
| Custom language models, speaker ID, PII redaction | Additional per-second charges |
Standard transcription bills a base per-second rate, rounded per request. Specialized variants like Transcribe Medical cost more, and features such as custom language models, speaker identification, and PII redaction add to the per-second rate. Rates fall in tiers at high volume, and a free tier covers initial usage.
What drives the bill
Total audio duration is the driver: a large library of long recordings costs in proportion to its hours. Premium variants and features multiply the base rate, so using Transcribe Medical or custom models where the standard tier would suffice over-pays. Transcribing audio you do not need (silence, irrelevant segments) also wastes seconds.
Controlling Transcribe cost
Transcribe only the audio you need, trimming silence and irrelevant segments, use the standard tier unless a specialized variant is genuinely required, enable premium features only where the use case needs them, and cache transcripts so you do not re-transcribe the same audio. As with any duration- billed AI service, sending less audio and using the right tier are the levers, mirroring Textract.
FAQ
How is AWS Transcribe priced?
Per second of audio transcribed, with a per-request minimum, at a base rate for standard transcription and higher rates for specialized variants like Transcribe Medical and features such as custom language models, speaker identification, and PII redaction. Rates fall in tiers at volume, with a free tier for initial usage.
How do I reduce Transcribe cost?
Transcribe only the audio you need by trimming silence and irrelevant segments, use the standard tier unless a specialized variant like Medical is genuinely required, enable premium features only where the use case needs them, and cache transcripts to avoid re-transcribing the same audio.
Why does Transcribe Medical cost more?
Because it is a specialized model tuned for medical terminology and accuracy requirements, priced at a higher per-second rate than standard transcription. Use it only for medical audio that needs its specialized recognition; for general audio, the standard tier is cheaper and sufficient.
What drives Transcribe cost?
Total audio duration, since billing is per second. A large library of long recordings costs in proportion to its hours. Premium variants and features multiply the base rate, and transcribing unnecessary audio (silence, irrelevant segments) wastes seconds. Duration and tier together set the bill.
Do Transcribe features like speaker identification cost extra?
Yes. Features such as custom language models, speaker identification, and PII redaction add to the per-second rate on top of base transcription. Enable them only where the use case genuinely needs them, rather than by default, to avoid inflating the per-second cost.
Does C3X estimate Transcribe cost?
Transcribe cost is driven by audio duration and tier, which are usage inputs. C3X prices the surrounding infrastructure, and you model your audio volume and the tier and features needed to estimate the per-second transcription charges.
What to do next
Estimate the infrastructure around your speech pipeline before you build it. C3X reads your Terraform and prices your resources against a live catalog. Start with the quickstart.
Share this post
Try C3X on your own Terraform
Free and open source. No API key required. One command to install, one command to estimate.