Velma API by Modulate

Audio-native API that reads emotion, intent and fraud risk from voice

Advanced API
Screenshot of Velma API by Modulate, Audio-native API that reads emotion, intent and fraud risk from voice

What is Velma API by Modulate?

Velma API by Modulate analyzes raw call audio to return emotions, speaker roles, 150+ behaviors, a diarized transcript and deepfake flags in one JSON response, with pricing from $0.75 per hour.

Velma API by Modulate is a voice analytics API that analyzes the audio of a conversation itself instead of working from a transcript alone. Its pitch is that ordinary speech-to-text followed by a language model throws away tone, hesitation, stress, pacing and speaker dynamics, which are often the signals that show what a call really means. Velma is built on an Ensemble Listening Model, described as more than 100 specialized sub-models, each tuned to a particular signal or task. A single call returns structured JSON. The output includes a diarized transcript with timestamps, speaker roles, detected emotions (more than 20), topics and sentiment per speaker, a short conversation summary, accent detection, PII and PHI tags ready for redaction, and deepfake detection on the same request. The behavior layer is the main differentiator: 50 behaviors ship by default, with 100 more available as templates, covering fraud, churn, compliance violations, harassment and escalation. Teams can also describe their own behaviors in plain language or upload SOPs, compliance documents and playbooks to steer what the model flags. Each detection comes with a confidence score and written reasoning. The target audience is developers and product teams building on voice: contact center and compliance tooling, fraud screening, live agent assist, voice agent guardrails, emotion-aware apps and conversation analytics. Audio can arrive from live streams, recordings, telephony or SIP, browser and mobile clients, and the service is reached over REST and WebSocket, so it can sit in front of an existing stack or replace the STT layer. Among alternatives, the comparison is with assembling a transcription service, a separate speech emotion model, a deepfake detector and prompt-engineered LLM logic. Velma packages these into one endpoint and advertises pricing from $0.75 per hour against a quoted $2.50 to $10 per hour for the pipeline approach. Its claims about deepfake accuracy and benchmark ranking come from the vendor, so buyers should test on their own audio.

How do you use Velma API by Modulate?

  1. 1Request API access
    Submit the signup request form on the Modulate platform to get free API access and credentials for Velma.
    Velma API by Modulate — Request API access
  2. 2Read the API documentation
    Open the docs to review the endpoints, REST and WebSocket options, accepted audio inputs and the structure of the JSON response.
    Velma API by Modulate — Read the API documentation
  3. 3Send your first audio file
    Submit a recording or live stream from your telephony, voice agent or app and inspect the diarized transcript, emotions, summary and detected behaviors.
  4. 4Define custom behaviors
    Describe the risks that matter to your business in plain language, or upload SOPs and compliance documents, then check the results against sample calls.
  5. 5Wire outputs into your workflow
    Route alerts, agent assist prompts, claims review or coaching flags from the structured output, and use PII tags for redaction.

Pros and cons

Pros

  • Single call returns transcript, emotions, behaviors, summary, accents and deepfake detection, so there is no multi-vendor pipeline to assembleAI
  • Analyzes the audio signal itself, capturing tone, stress and prosody that transcript-only approaches discardAI
  • Custom behaviors can be written in plain language or guided by uploaded SOPs and playbooks, with no fine-tuningAI
  • Supports both REST and WebSocket, covering recorded files and real-time streamsAI
  • Published entry price of $0.75 per hour is well below the pipeline cost range the vendor quotesAI

Cons

  • Accuracy, ranking and cost-saving claims are vendor-stated, and independent validation on a buyer's own audio is neededAI
  • Voice-based emotion and deception inference can raise privacy, consent and regulatory questions for call monitoringAI
  • Access starts with a signup request rather than instant self-serve, so quick experiments may be slowerAI
  • Only the entry rate is public on the homepage, so costs at volume or with extra features are unclearAI
  • Built for developers integrating an API, with no stated ready-made dashboard for non-technical teamsAI

How much does Velma API by Modulate cost?

Free trial

Free API access can be requested through the platform signup form. Limits and duration are not stated on the homepage.

Learn more

Pricing

Starting at $0.75 per hour of audio, compared by the vendor with $2.50 to $10 per hour for a transcription plus LLM pipeline. Free API access is available by request. Higher-volume terms are not listed.

Learn more

Support

A documentation site is available for developers. Dedicated support channels and service levels are not described on the homepage.

Learn more

Integrations

Accepts audio from live streams, recordings, telephony and SIP, voice agents, and browser or mobile clients. Delivered through REST and WebSocket APIs for real-time alerts and agent assist.

Learn more

Features

Diarized transcript, 20+ emotions, speaker identification, 150+ behaviors with reasoning, custom behaviors in plain language, topics and sentiment, summaries, accent detection, deepfake detection and PII/PHI tagging in one API call.

Learn more

Frequently asked questions about Velma API by Modulate

  • How much does Velma API by Modulate cost?
    Starting at $0.75 per hour of audio, compared by the vendor with $2.50 to $10 per hour for a transcription plus LLM pipeline. Free API access is available by request. Higher-volume terms are not listed.
  • Does Velma API by Modulate offer a free trial?
    Free API access can be requested through the platform signup form. Limits and duration are not stated on the homepage.
  • Does Velma API by Modulate have an API?
    Yes, Velma API by Modulate offers an API.
  • How do you use Velma API by Modulate?
    The walkthrough on this page covers 5 steps: 1. Request API access 2. Read the API documentation 3. Send your first audio file 4. Define custom behaviors 5. Wire outputs into your workflow.
  • What platforms does Velma API by Modulate support?
    Velma API by Modulate is available on Web App.
  • What does Velma API by Modulate integrate with?
    Accepts audio from live streams, recordings, telephony and SIP, voice agents, and browser or mobile clients. Delivered through REST and WebSocket APIs for real-time alerts and agent assist.

Status

StatusActive
Views0
Outbound clicks0
Added10/8/2026

Platforms

Web App

Pricing

PaidFree Trial

Categories