Free trial
Free API access can be requested through the platform signup form. Limits and duration are not stated on the homepage.
Learn moreAudio-native API that reads emotion, intent and fraud risk from voice

Velma API by Modulate analyzes raw call audio to return emotions, speaker roles, 150+ behaviors, a diarized transcript and deepfake flags in one JSON response, with pricing from $0.75 per hour.
Velma API by Modulate is a voice analytics API that analyzes the audio of a conversation itself instead of working from a transcript alone. Its pitch is that ordinary speech-to-text followed by a language model throws away tone, hesitation, stress, pacing and speaker dynamics, which are often the signals that show what a call really means. Velma is built on an Ensemble Listening Model, described as more than 100 specialized sub-models, each tuned to a particular signal or task. A single call returns structured JSON. The output includes a diarized transcript with timestamps, speaker roles, detected emotions (more than 20), topics and sentiment per speaker, a short conversation summary, accent detection, PII and PHI tags ready for redaction, and deepfake detection on the same request. The behavior layer is the main differentiator: 50 behaviors ship by default, with 100 more available as templates, covering fraud, churn, compliance violations, harassment and escalation. Teams can also describe their own behaviors in plain language or upload SOPs, compliance documents and playbooks to steer what the model flags. Each detection comes with a confidence score and written reasoning. The target audience is developers and product teams building on voice: contact center and compliance tooling, fraud screening, live agent assist, voice agent guardrails, emotion-aware apps and conversation analytics. Audio can arrive from live streams, recordings, telephony or SIP, browser and mobile clients, and the service is reached over REST and WebSocket, so it can sit in front of an existing stack or replace the STT layer. Among alternatives, the comparison is with assembling a transcription service, a separate speech emotion model, a deepfake detector and prompt-engineered LLM logic. Velma packages these into one endpoint and advertises pricing from $0.75 per hour against a quoted $2.50 to $10 per hour for the pipeline approach. Its claims about deepfake accuracy and benchmark ranking come from the vendor, so buyers should test on their own audio.


Free API access can be requested through the platform signup form. Limits and duration are not stated on the homepage.
Learn moreStarting at $0.75 per hour of audio, compared by the vendor with $2.50 to $10 per hour for a transcription plus LLM pipeline. Free API access is available by request. Higher-volume terms are not listed.
Learn moreA documentation site is available for developers. Dedicated support channels and service levels are not described on the homepage.
Learn moreAccepts audio from live streams, recordings, telephony and SIP, voice agents, and browser or mobile clients. Delivered through REST and WebSocket APIs for real-time alerts and agent assist.
Learn moreDiarized transcript, 20+ emotions, speaker identification, 150+ behaviors with reasoning, custom behaviors in plain language, topics and sentiment, summaries, accent detection, deepfake detection and PII/PHI tagging in one API call.
Learn more




