Unreal Speech

A low-cost text-to-speech API for streaming, long-form audio, and word timestamps.

Intermediate API
Screenshot of Unreal Speech, A low-cost text-to-speech API for streaming, long-form audio, and word timestamps

What is Unreal Speech?

Unreal Speech is a developer-focused text-to-speech API offering low-latency streaming, 48 voices across multiple languages, per-word timestamps, and pricing built around high-volume production use rather than casual narration.

Unreal Speech is a text-to-speech API aimed at developers who need to turn written text into spoken audio programmatically rather than through a consumer-facing app. It converts text via REST and websocket endpoints, with streaming synthesis starting in roughly 300 milliseconds and support for single requests generating up to 10 hours of audio. Voice output can include per-word or per-sentence timestamp data, which is useful for building synchronized captions, karaoke-style text highlighting, or accessibility features tied to playback position. The tool is built for engineering teams shipping audio at volume: audiobook and e-reading platforms, news and blog readers, IVR and phone systems, and e-learning tools that need narrated content generated automatically. A customer testimonial from the CEO of Listening.com describes processing over 10,000 pages per hour through the service, which signals the product is oriented toward production throughput rather than one-off narration. In practice, integration happens through one of four endpoints chosen by text length and latency needs: /stream for short clips returned almost instantly, /speech for medium-length text with timestamp URLs, /synthesisTasks for asynchronous jobs up to 500,000 characters, and a /streamWithTimestamps websocket for real-time word-level timing during playback. Developers select from 48 voices across up to nine languages, and can adjust speed, pitch, bitrate, and output format, including standard MP3, phone-call optimized audio, and PCM mu-law. Official SDKs cover Python, Node.js, and React Native, with sample code provided for each endpoint and a GitHub repository demonstrating word-highlighting playback. Unreal Speech positions itself squarely on cost and throughput against Amazon Polly, Microsoft, Google, ElevenLabs, and Play.ht, publishing its own price and quality comparisons across fiction, non-fiction, and conversational text. For teams already invested in a premium voice-cloning service for highly expressive narration, this tool reads more as a cost-optimized production alternative than a creative playground; a 250,000-character free allotment and a live demo let teams sample voice quality before adopting it for larger workloads.

How do you use Unreal Speech?

  1. 1Claim a free API key
    Sign up on the Unreal Speech site to get an API key and a 250,000-character free allotment before writing any integration code.
  2. 2Preview voices in the live demo
    Use the on-site demo to test different content types, voices, languages, and audio formats to shortlist a voice before building against the API.
  3. 3Pick the right endpoint for your text length
    Use /stream for short instant clips, /speech for up to 3,000 characters with timestamps, /synthesisTasks for asynchronous jobs up to 500,000 characters, or the /streamWithTimestamps websocket for real-time word timing.
  4. 4Send your first request
    Call the chosen endpoint with your text, VoiceId, bitrate, speed, and pitch parameters using the provided Python, Node.js, React Native, or Bash sample code.
  5. 5Retrieve audio and timestamp data
    Save the returned MP3 or PCM audio and, if requested, parse the TimestampsUri JSON to sync word or sentence timing with playback.
  6. 6Monitor usage as volume grows
    Track consumption against the pay-as-you-go rate of $8 per million characters, or move to the enterprise plan once monthly volume approaches its included character allotment.

Pros and cons

Pros

  • Published price comparisons position it well below Amazon, Microsoft, Google, ElevenLabs, and Play.ht for equivalent volumeAI
  • Streaming latency of around 300 milliseconds suits real-time or near-real-time playback scenariosAI
  • Single requests can generate up to 10 hours of audio, useful for full audiobook or long-document narrationAI
  • Per-word and per-sentence timestamp data supports synchronized captioning and text-highlighting featuresAI
  • A 250,000-character free allotment lets developers evaluate voice quality before payingAI

Cons

  • Purely an API product with no dashboard-driven workflow for non-technical users beyond the demo pageAI
  • 48 voices across roughly nine languages is a narrower catalog than some multilingual TTS competitorsAI
  • The enterprise tier jumps to a $4,999 monthly commitment, a steep step up from the free allotment for growing usageAI
  • Quality and price comparisons on the homepage are self-published using public competitor pricing, so real-world costs on custom plans may differAI
  • The live demo restricts text input to 250 characters, limiting how thoroughly voice quality can be judged before integrating the APIAI

How much does Unreal Speech cost?

Free trial

New users get 250,000 characters free after signing up for an API key, enough to test voices, languages, and endpoints before committing to paid usage.

Learn more

Pricing

Includes a free tier of 250,000 characters with a free API key. Beyond that, usage is billed at $8 per 1 million additional characters. An enterprise plan is available at $4,999 per month, which includes 625 million characters (an estimated ~14,000 hours of audio); higher-volume needs are handled through custom inquiry.

Learn more

Support

Support appears centered on API documentation and sample code, with a contact channel for high-volume or custom enterprise inquiries; no live chat or phone support is mentioned on the homepage.

Integrations

Accessible via REST endpoints and a websocket streaming endpoint, with official SDKs for Python, Node.js, and React Native, plus sample code and a GitHub demo repository for word-highlighting playback.

Features

Streaming synthesis in about 300ms, requests generating up to 10 hours of audio, per-word and per-sentence timestamps, 48 voices across up to nine languages, adjustable speed/pitch/bitrate, and multiple output formats including MP3, phone-call audio, and PCM mu-law.

Frequently asked questions about Unreal Speech

  • How much does Unreal Speech cost?
    Includes a free tier of 250,000 characters with a free API key. Beyond that, usage is billed at $8 per 1 million additional characters. An enterprise plan is available at $4,999 per month, which includes 625 million characters (an estimated ~14,000 hours of audio); higher-volume needs are handled through custom inquiry.
  • Does Unreal Speech offer a free trial?
    New users get 250,000 characters free after signing up for an API key, enough to test voices, languages, and endpoints before committing to paid usage.
  • Does Unreal Speech have an API?
    Yes, Unreal Speech offers an API.
  • How do you use Unreal Speech?
    The walkthrough on this page covers 6 steps: 1. Claim a free API key 2. Preview voices in the live demo 3. Pick the right endpoint for your text length 4. Send your first request 5. Retrieve audio and timestamp data 6. Monitor usage as volume grows.
  • What platforms does Unreal Speech support?
    Unreal Speech is available on Web App.
  • What does Unreal Speech integrate with?
    Accessible via REST endpoints and a websocket streaming endpoint, with official SDKs for Python, Node.js, and React Native, plus sample code and a GitHub demo repository for word-highlighting playback.

Status

StatusActive
Views0
Outbound clicks0
Added8/4/2026

Platforms

Web App

Pricing

FreemiumPaidSubscription