Skip to contentx402 is live on Robinhood Chain: agents can now pay for any model, per call, in USDG →
OmniRail
← All models

Microsoft AI · Voice model

MAI-Voice-2-Flash

microsoft/mai-voice-2-flash
  • Voice
Released Jul 23, 2026

MAI-Voice-2-Flash is a low-latency text-to-speech model from Microsoft AI for voice agents, assistants, call centers, accessibility, narration, and other interactive applications. It generates expressive 24 kHz mono speech across 15...

Specifications

Input
text, up to 20,000 characters per call
Output
mp3 (default) or raw pcm
Voice
provider voice ids via `voice`; omitted = the provider's default when it has one
Billing
per character of input text

Use it

POST /api/v1/audio/speech · MCP speak · full reference

curl https://omnirail.org/api/v1/audio/speech \
  -H "Authorization: Bearer $OMNIRAIL_KEY" -H "Content-Type: application/json" \
  -d '{"model":"microsoft/mai-voice-2-flash","input":"Hello from OmniRail. Every model, one key.","response_format":"mp3"}' \
  --output hello.mp3