Remy Voice

Voice agents that
show their work.

Most voice agents are a black box: you hear an answer and take it on faith. With Remy, you build on a best-in-class voice pipeline without wiring it together yourself.

A demo built with Remy. No sign-up.
The platform

Every part best-in-class. All of it wired for you.

Build with one realtime model end to end, or piece together a best-in-class model at every stage. Swap any part, any time.

Speech-to-speech
One realtime model hears and speaks. Lowest latency, the way Cove runs below.
GPT RealtimeGemini LiveGrok Voice
Cascading
Compose a best-in-class model at every stage. Swap any part anytime.
Listen
Speech to text
DeepgramWhisper
Think + tools
Reasoning and tool calls
OpenAIGeminiClaude
Speak
Text to speech
CartesiaElevenLabs
Transport & telephony
Every call arrives and leaves over infrastructure we run, so you never stand up a media server or a phone carrier yourself.
LiveKitTelnyx
Capabilities
What an agent does once it's on the line.
Tool-calling on real methodsCalls your backend and acts on what comes back.
Vector search with citationsAnswers from your docs, every claim cited.
In-call verificationConfirms identity with an SMS or email code.
Instant barge-inCut in any time and it stops talking to listen.
Session contextStarts the call already knowing the caller.
Inbound & outbound callsIn the browser, or on a real phone number.

Plus live captions, results that render on screen, and swapping the model or voice mid-call.

Standalone speech
Text-to-speech and transcription, ready for any app, agent or not.

157 voices · 20 speech models · metered per minute or per million characters

Hear them all

Talk to Cove.

Cove is a demo, not a product: a made-up smart-home audio brand we built with Remy to show what the platform can do. Everything it does is real, though. Talk to it like a customer, and once you verify your number by text, it pulls up your account and takes real action, like rescheduling a technician visit or placing a callback.

Pipeline
Speech-to-speech
One model listens, thinks, and speaks.
Model
gpt-realtime-2.1
Marin voice
Transport
LiveKitBrowser · WebRTC
TelnyxPhone · PSTN
Toolbox
Look up policies
Load account
Check order status
Run diagnostics
Reschedule a visit
Request a callback
Cove Support
Online
Ready

Hi, I'm Cove's assistant.

Ask about an order, how something works, or your account.

Try asking
Account actions verify your number by text first.
What Cove surfaces
Answers, your account, and confirmations appear here
Citations
Activity
session · —00:00
StandbyAwaiting first tool call

Or call it for real.

The same agent you just talked to, now live on a real phone call.

US number · standard call rates apply.

Rate card

Priced by the unit, metered to the second.

Every voice model on the platform, at its real per-unit rate. Compose any pipeline and pay for exactly what runs, billed per second of audio and per token or character of text. No bundles, no seat minimums.

26 voice models · four layers · pass-through metering · no minimums
Realtime

One model listens, thinks, and speaks. Billed per minute of conversation.

ProviderModelRate
OpenAI
gpt-realtime-2.1in this demo~$0.30/min
audio $32 / $64 · text $4 / $24 · cached $0.40 · per 1M tok
OpenAI
gpt-realtime-2.1-mini~$0.05/min
audio $10 / $20 · text $0.60 / $2.40 · cached $0.06 · per 1M tok
xAI
grok-voice-think-fast-2.0$0.08/min
text input $4 · per 1M tok
Google
gemini-2.5-flash-native-audio~$0.02/min
audio $3 / $12 · text $0.50 / $2 · per 1M tok
Google
gemini-3.1-flash-live~$0.02/min
audio $3 / $12 · text $0.75 / $4.50 · per 1M tok
Speech-to-text

Transcription for composed pipelines. Billed per minute of audio.

ProviderModelRate
Deepgram
Nova-3$0.0058/min
OpenAI
Whisper-1$0.006/min
ElevenLabs
Scribe v2$0.0067/min
ElevenLabs
Scribe v1$0.0067/min
DeepInfra
Whisper v3 Turbo$0.0002/min
DeepInfra
Whisper v3$0.00045/min
DeepInfra
Voxtral Mini 3B$0.001/min
DeepInfra
Qwen3 ASR 1.7B$0.00045/min
OpenAI
GPT-4o Transcribe$2.50 / $10per 1M tok
Text-to-speech

Voices for composed pipelines. Billed per character or per token.

ProviderModelRate
OpenAI
tts-1$15/1M chars
OpenAI
tts-1-hd$30/1M chars
ElevenLabs
Turbo v2.5$90/1M chars
ElevenLabs
Turbo v2$90/1M chars
ElevenLabs
Multilingual v2$180/1M chars
ElevenLabs
v3$180/1M chars
Cartesia
Sonic-3$37/1M chars
Qwen
Audio 3.0 TTS Plus$27.59/1M chars
DeepInfra
Qwen3 TTS$20/1M chars
Google
Gemini 3.1 Flash TTS$1 / $20per 1M tok
OpenAI
GPT-4o Mini TTS$0.60 / $2.40per 1M tok
MiniMax
Speech 2.8 HD$0.007/run
Telephony & infrastructure

Transport, carriage, and numbers. The managed layer both architectures run on.

ProviderModelRate
Remy
Browser transport (WebRTC)$0.00
no per-minute transport fee
Remy
Phone carriage (PSTN / SIP)$0.006/min
Remy
Dedicated phone number$1.00/month

List prices in USD. Realtime rates are approximate: the per-token rate under each model is what actually meters usage. Your bill is the sum of exactly what each call uses. Volume and committed-use pricing available.