Skip to contentx402 is live on Robinhood Chain: agents can now pay for any model, per call, in USDG →
OmniRail
← All models

Microsoft AI · Voice model

MAI-Voice-2.1

microsoft/mai-voice-2.1
  • Voice
Released Oct 1, 2026

MAI-Voice-2.1 is Microsoft AI's highest-fidelity, most expressive text-to-speech model. It produces natural, studio-grade speech across 23 languages, with detailed prosody, nuanced expressiveness, and speaker consistency over long-form content. It is...

Specifications

Input
text, up to 20,000 characters per call
Output
mp3 (default) or raw pcm
Voice
provider voice ids via `voice`; omitted = the provider's default when it has one
Billing
per character of input text

Use it

POST /api/v1/audio/speech · MCP speak · full reference

curl https://omnirail.org/api/v1/audio/speech \
  -H "Authorization: Bearer $OMNIRAIL_KEY" -H "Content-Type: application/json" \
  -d '{"model":"microsoft/mai-voice-2.1","input":"Hello from OmniRail. Every model, one key.","response_format":"mp3"}' \
  --output hello.mp3